Navigation
Knowledge record format
Knowledge record format: what a record consists of, how it is stored next to the code and how the assistant reads it
Knowledge records live as files in the .gitriver/knowledge/ directory of
the main repository — not in a separate database and not in a wiki. That’s
why a record gets a branch, review, CODEOWNERS, branch protection, and
revert, and travels in the same branch as the code change that produced it:
one review, one approval, one revert.
One file, one record. Two people editing different records are editing different files, so there are no merge conflicts between them.
File structure
A Markdown body under a YAML header (frontmatter) between --- lines:
---
id: 0192ec1a-7b1a-7def-abcd-0987654321fe
origin: agent
contract:
- mechanism: on_query
params:
words: ["deploy", "rollback"]
- mechanism: on_tool
params:
name: run_migration
links:
- 0192ec1b-0000-7000-8000-000000000001
---
The first complete thought in the record — this is also the excerpt shown in lists.
Next comes the reasoning: why this is the way to do it, what to try, when it doesn't work.
Header fields
| Field | Required | Meaning |
|---|---|---|
id |
yes | Stable record identifier, UUIDv7. Lives inside the file: survives renaming the file and any edit to the text. Impression and read counters are tracked against it |
contract |
no | Retrieval contract: an array of declarations {"mechanism": ..., "params": {...}}. Empty means “never surface” — such a record is considered unreachable but is still stored |
origin |
no | Provenance: human or agent. Defaults to agent. The field in the file is the client’s claim, not proof: the guarantee of provenance is the pull request’s approval, not the header |
links |
no | Links to other records — a list of their stable identifiers |
summary |
no | An explicit display excerpt. See the override rule below |
File names are human-readable and free to change — the identifier doesn’t
live in the path and so it’s never lost. Records the server assembles
itself (see “How knowledge gets out”) are placed under the name <id>.md;
such a file can be freely renamed afterward.
Body and display excerpt
The excerpt shown in lists is derived from the body: the first paragraph up to a blank line, capped at 512 bytes at a character boundary.
summary overrides the derived excerpt, but it must say something
different than the body. An override that repeats the derived excerpt
(after whitespace normalization) is flagged in the parse diagnostics, and
the derived excerpt is used instead. A derived excerpt that’s meaningless
or runs into the cap also shows up in the parse diagnostics — fix it in
the record itself.
Strict on write, tolerant on read
The rule is decided by where the data came from:
- creating a record through the API — strict parsing: an unknown key is rejected on the spot as a typo;
- parsing a file from the repository — tolerant: the file may have been written by a tool of a newer version than the server.
Tolerance distinguishes not understood from wrong:
- not understood — a key or mechanism this version doesn’t recognize. The record is accepted in full, declared mechanisms are stored as written, and the record counts as unreachable until the server learns those mechanisms; once the server is updated the record comes back to life on its own, with no need to rebuild the index. An unknown key is named together with the closest known one: “the key “on_trigers” is unknown, the closest known one is “contract”” — a typo from someone editing by hand; “nothing similar exists” — the record is newer than the server’s version;
- wrong — an invalid value for a known field (
idisn’t a UUIDv7), a file larger than the limit, a duplicate identifier within a batch, an empty body. Such a record is skipped, and the rest keep getting parsed: one broken file doesn’t affect the others.
Unrecognized mechanisms are rolled up in the parse report: “3 records require the mechanism “on_path”, which this version does not support.” That line tells an administrator it’s time to update the server.
The parse report reaches the author in response to the push — as remote:
lines next to the git push output (over HTTP and over the built-in SSH
server): one line per remark about a file, “path: reason”, and one summary
line per unrecognized mechanism. The push itself goes through.
Limits
| Limit | Value |
|---|---|
| Size of a single file | 64 KB |
| Total size of a parse batch | 10 MB |
Exceeding a limit is a diagnostic and the file gets skipped, not a rejected push.
Both limits apply on input too. A draft is checked against the size of the file it will become, header included, not just its text: an accepted draft can always be delivered. A batch gets cut by the parse limit when it’s assembled: the remainder goes out in the next batch instead of being rejected.
Parsing on push
The knowledge index is filled from two sources: drafts are written
straight into it, and .gitriver/knowledge/ files are parsed when
changes are received — only the files changed in that push, with content
read from the new commit.
- Parsing knowledge never fails a push: any error becomes a log entry.
- File-based records reach the index only from the default branch. Knowledge from a branch that hasn’t been merged yet is visible to the team as drafts.
- Parsing still runs on every push: the author learns about a typo in a record right away, not after the merge.
- Merging a pull request through the interface or a merge queue is parsed too: knowledge becomes available as soon as the request is merged, without waiting for the next push to the default branch.
- Deleting a file removes the index record only on the default branch. If the index momentarily drifts from the tree, the drift resolves itself on the next read, with no action from the administrator.
- Reading knowledge always goes through the index, not a branch’s files: one branch’s drafts are visible from other branches even before the merge. Results are filtered by the caller’s permissions on each repository.
How knowledge gets out
An agent doesn’t commit every thought. It writes a draft — a draft is immediately visible to the whole team with access to the repository and is marked unverified — and knowledge gets out in a batch: every draft accumulated on one branch becomes files in a single commit and goes through a single review.
There are two paths, and the first is the main one.
- Into the same branch as the code. Files are placed on the branch the agent was working on: the knowledge arrives inside the developer’s own pull request — one review, one approval, no separate step at all.
- On a separate branch with its own pull request. The branch is
named
knowledge/<source branch>-<tag>. A fallback path: needed when the source branch is already gone (merged and deleted), or when the knowledge isn’t tied to a code change.
Knowledge has no approval flow of its own on either path: it gets out through an ordinary pull request, and branch protection, CODEOWNERS, required approvals, and merge checks apply to it exactly as they do to any other request. The server’s own file-writing is subject to branch protection the same as a push: a batch won’t land on a branch where pushes are forbidden.
The server sets the request’s provenance
A pull request opened by the knowledge-assembly process is marked with machine provenance. The server sets the mark: there is no such field in the request body for creating a pull request, and a client cannot declare its own request machine-generated. This is the same rule that governs machine-generated pull request comments.
What happens to a draft
Nothing — until approval. A draft is marked delivered (where, when, which pull request it’s waiting in) and doesn’t get picked up into a batch a second time, but it keeps existing as a draft: if the request is rejected and the branch deleted, the knowledge isn’t lost.
What removes a draft is its file landing on the default branch. If a batch went straight to the default branch (a small team working without branches), this happens immediately, and no pull request appears at all. The index record stays the same one throughout — the identifier is shared, so impression and read counters survive the draft turning into a file.
Interface
Knowledge is read and written through /api/v1; reads always go through
the index, never through the tree’s files.
| Call | What it does |
|---|---|
GET /repos/{owner}/{name}/knowledge |
Repository index: {items, partial?} — records with their contract and a reachability flag |
GET /repos/{owner}/{name}/knowledge/drafts |
Repository drafts — all of them, not just your own; a page {items, next_cursor} (per_page, after) |
POST /repos/{owner}/{name}/knowledge/drafts |
Write a draft |
DELETE /repos/{owner}/{name}/knowledge/drafts/{id} |
Remove a draft |
POST /repos/{owner}/{name}/knowledge/drafts/deliver |
Assemble the drafts accumulated on the branch into files |
POST /repos/{owner}/{name}/knowledge/shown |
Mark a batch of records as shown |
POST /repos/{owner}/{name}/knowledge/entries/{id}/read |
Mark that a record was opened |
GET /repos/{owner}/{name}/knowledge/stats/me |
My own impression and read counters: {items, partial?} |
GET /repos/{owner}/{name}/knowledge/stats/team |
Team roll-up, the same shape (Pro). There’s a row for EVERY record in the index, including ones never shown — with zeros |
GET /repos/{owner}/{name}/knowledge/stats/dead |
Dead-memory report (Pro): {items, partial?, min_shown, min_age_days}. Each row has an ailment — shown_not_read, unreachable (with unmet_mechanisms), or never_shown. Thresholds are set by min_shown and min_age_days, and the ones actually applied are returned in the response |
GET·PUT·DELETE /admin/knowledge/limits/{users|groups}/{id} |
Draft limit per owner: viewing and clearing in any edition, setting and changing (PUT) in the Max edition |
Displays are marked in a batch — up to 256 records per call.
Displays can be marked as often as they actually happen: this call’s rate limit (600 requests per minute per token) has plenty of headroom. Reading the counters is always exact, even right after a display is marked.
The index is returned in full, with no pagination — matching against the
contract is done by the consumer. A protective cap of a thousand records
applies; if the result got cut, the response carries the
partial: "truncated" caveat. Counter summaries are bounded by the same
cap and the same caveat. The absence of the partial field means a
complete response, so a client that doesn’t know about it keeps working
with no changes needed.
The index doesn’t carry a record’s body — for a file-based record it’s
read the ordinary way, as a file, via file_path; for a draft it comes
back in the list of drafts.
The permissions are knowledge:read and knowledge:write; they’re
derived from the repository access level (reading knowledge = reading the
repository, writing = writing). A token is always required, even on a
public repository.
The edition split is expressed through ADDRESSES, not a request field: one repository’s knowledge and personal counters work fully in Community; only the team roll-up and the dead-memory report (Pro) and setting draft limits (Max) are paid. A paid call is either available in full or rejected in full — there’s no way to get the paid part of a response through the free call. More detail — licensing.md.
A license rejection carries the machine code license_required, distinct
from a permissions rejection: the client doesn’t need to parse translated
text.
Example of a complete file
---
id: 0192ec1a-7b1a-7def-abcd-0987654321fe
origin: human
contract:
- mechanism: on_query
params:
words: ["migration", "sqlx"]
summary: Preparing the database before sqlx migrate — otherwise migrations fail on permissions
---
Sqlx migrations need a schema owner, not a superuser.
A migration applied as superuser creates objects owned by that superuser, and
the next migration run under an ordinary account fails on permissions. Set up
a separate schema-owner role and apply migrations under it.