vibecommitdocsinstallpricinggithubsign instart →
docs · disclosure

What we upload

VibeCommit records the coding-agent sessions you run in a repository you turn capture on for. This page is the long version of the disclosure vibecommit connect prints before it asks whether to turn capture on for a repository: what leaves your machine, what is removed before the record is written, and what can be read back out afterwards.

What leaves your machine

Capture is per repository and off until you turn it on. Once it is on, hooks inside your coding agent upload that repository's session transcript as you work — on each assistant turn, before the agent compacts its history, and at the end of the session. Nothing waits for a commit, and nothing asks a second time.

Uploaded, in full:

  • Your prompts — the whole text, not a summary.
  • The agent's replies — likewise.
  • The arguments to every tool call — including the file contents inside Read, Edit, MultiEdit, Write and NotebookEdit records, for files inside the repository you consented to. Content stored beside a path outside it is replaced before upload — see Files outside the repository below.
  • The output of the commands the agent ran.
  • File paths, including paths inside private repositories.
  • Repository metadata — the commit SHA, the branch, the remote, the commit message, and the diff.
  • Sub-agent transcripts. Work your agent delegates runs in its own transcript file, and those files are captured alongside the main one.

Files outside the repository

The consent unit is the repository; the capture unit is the session, and a session can touch more than the repository it happens to be sitting in. Some of that out-of-tree content is held back before it leaves your machine and some of it is not, and the line between them is worth stating exactly, because it is a rule about where a file sits and not about what is in it.

Structured file reads and writes are filtered by path. When a tool record names a file — the file_path, filePath, notebook_path or notebookPath key of a Read, Edit, MultiEdit, Write or NotebookEdit — and that file resolves outside the repository you consented to, the content stored beside that path is replaced on your machine, before anything is uploaded. What reaches us is a marker of the same byte length:

[OUT-OF-TREE: 58 bytes withheld..........................]

The count is the size of the content that stayed behind. The dots are padding, and they are load-bearing rather than decorative: the client uploads byte ranges and our ingest advances its position by the length of what arrived, so a replacement that changed length would put those two counts out of step and every later upload would be recorded as a gap — a capture-health surface reporting data loss that is not happening. A span too short for the counted form gets the bare [OUT-OF-TREE] instead. The rule holds for the resolution step too: a path the client cannot resolve is treated as outside, so its content is replaced rather than uploaded on the assumption it was yours.

Command output is not filtered by path. A shell command names no file in a key the client reads, so if the agent runs a command that prints a file from outside the repository, that output is part of the session and is uploaded with it. This is the same sentence this page carried before the filter existed, and it is still true.

The filter is a path comparison, not a content scan, and it is deliberately partial. Alongside command output, it also leaves:

  • File paths themselves. A path is what triggers the comparison; it is never what gets replaced. An out-of-tree path is uploaded in full, as are paths inside private repositories.
  • File content under a key nobody enumerated. Ten content keys are read, plus the changed lines inside a structuredPatch. A tool that carries file content under some other name is uploaded as it was written.
  • Values too short to hold a marker. Because the marker has to match the byte length it replaces, a value shorter than [OUT-OF-TREE] is left as it was and counted as skipped — a short leak in preference to broken accounting.
  • Anything on a line the client cannot rewrite byte for byte. A line that does not survive the client's own round-trip check is returned exactly as it was found, unredacted, rather than guessed at; so is an upload whose edge falls inside a multi-byte character. The filter reverts rather than risk mangling a line, because a mangled line would be permanent.

What the scrubber removes

Before a turn is written to the record, a scrubber replaces matched secrets with a [REDACTED:<kind>] token. The token stays in the text, so you can see that something was removed and what kind of thing it was, without the value being there.

Twelve patterns, eleven tokens

The scrubber runs twelve patterns and emits eleven distinct tokens. The two numbers differ for exactly one reason: env is matched by two patterns — a KEY=value line on its own, and an inline assignment in code — and both produce [REDACTED:env]. The list below is keyed by token, because the token is what you will actually see in your own record, so it has eleven rows and the env row names both of its patterns.

  • [REDACTED:env]
    Environment-variable values. Two patterns: a KEY=value line on its own, and an inline assignment — process.env.X = "…" or os.environ["X"] = "…". Both produce the same token.
  • [REDACTED:jwt]
    JSON Web Tokens — three base64url segments where the header and payload both begin eyJ.
  • [REDACTED:aws_key]
    AWS access key IDs beginning AKIA. Short-term ASIA session keys are not matched.
  • [REDACTED:aws_secret]
    AWS secret access keys — a 40-character token with aws_secret_access_key within 200 characters of it. Without that cue nearby, it is left alone.
  • [REDACTED:gcp_key]
    Google Cloud service-account JSON containing both "type": "service_account" and a BEGIN PRIVATE KEY block.
  • [REDACTED:azure_key]
    Azure storage account keys — 88 characters of base64 ending ==, with azure, storage or account key nearby.
  • [REDACTED:azure_conn]
    The AccountKey= segment of an Azure storage connection string.
  • [REDACTED:stripe]
    Stripe keys — sk_, rk_ and pk_, in both live and test form.
  • [REDACTED:github_pat]
    GitHub tokens — ghp_, gho_, ghu_, ghs_ and ghr_.
  • [REDACTED:credit_card]
    Card numbers of 13 to 19 digits that pass the Luhn checksum. A digit run that fails Luhn is left alone, and the published test numbers (4242… and the rest) are checked against an exact list and left alone too.
  • [REDACTED:ssn]
    US Social Security numbers written XXX-XX-XXXX, gated on the SSA allocation rules — area not 000, 666 or 900–999; group not 00; serial not 0000. A formatted number failing those rules is left alone.

Nine of those are secret kinds. Two — credit_card and ssn — are structured personal data, and both are gated: a match is only redacted if it also passes a validity check, so a random digit run stays as it was written.

What the scrubber does not do

It is a pattern list, not a general secret detector. Three limits are worth stating outright:

  • A secret with no pattern is uploaded. There is no catch-all. A generic high-entropy detector exists as a configuration hook in the scrubber, but the pattern itself does not run in this release — so a credential that does not look like one of the eleven kinds above travels with the rest of the session and is stored with it.
  • Bare nine-digit SSNs are not matched. Only the formatted XXX-XX-XXXX shape is. An unseparated nine-digit run collides with order numbers and record IDs too often to match on shape alone.
  • Names, street addresses and organisation names are not matched at all. Recognising those takes entity recognition rather than pattern matching, and that is not in this release.

The scrubber never records a matched value

This one is a property of the code rather than a promise about our conduct, which is why it is worth stating precisely. The scrub function returns two things: the redacted text, and a count per kind — env: 2, github_pat: 1. The matched value is not among what it returns, so the internal review that reads redaction distribution reads counts and has nothing else to read.

Where redaction runs, and why

The scrubber runs on our ingest servers, not on your machine. The transcript is uploaded first and scrubbed on arrival, before the turn is hashed and written to the record. Stated plainly, and it is the sentence this page exists to make sure you read: a secret that one of the twelve patterns matches still travels to us once, and is removed before the turn is hashed and stored. A secret no pattern matches is stored with the turn.

The reason is the shape of the record. A turn's identity is the SHA-256 of its scrubbed bytes, so the scrub has to be the same function for everyone — one implementation, one version, applied at one place. If every machine scrubbed locally with whatever version it happened to be running, two identical turns would hash differently and the record would stop addressing its own content. Running the scrubber once, on arrival, before hashing, is what keeps a hash a property of the content.

That same constraint is why the other filter runs in the opposite place. Whether a path falls outside the repository you consented to depends on where a file sits rather than what it says — a fact about your machine's layout, not about the bytes. It is also not a fact we could check if we wanted to: by the time a transcript reaches us, the machine that could answer the question is gone, and rewriting content after arrival would break the same hash-identity property described above. So it could only ever be applied before upload, and that is exactly where it is applied — on your machine, at the one point the client turns transcript bytes into a request body, before compression and before the upload.

The two are different mechanisms and it is worth not merging them. The scrubber runs on our servers, keys on what the bytes say, and leaves a [REDACTED:<kind>] token behind. The out-of-tree filter runs on your machine, keys on where the file sits, and leaves a padded [OUT-OF-TREE: N bytes withheld] marker behind. The client does not scrub secrets — it compares paths. A credential sitting in a file inside the repository you consented to is uploaded and then handled by the scrubber above, with the limits the scrubber section already states.

Sub-agent transcripts go through the same filter. Work your agent delegates is written to its own transcript file, and both the main transcript and each delegated one are read through one function and filtered against the same project root before upload — one reader, one filter, both kinds of file.

For org owners

Capture is consented twice. The developer consents on their own machine, per repository. The organisation that owns the repository's namespace consents separately, and that second gate is yours.

What pends, and what does not

An organisation's capture_policy defaults to approval_required. Under that default, already-approved members keep capturing; only a new member's first capture in an unclaimed namespace pends. An outstanding request holds that one developer's first capture in that one namespace. It does not stop the members who are already capturing, and it is not a switch that pauses the team.

What captures anyway, while a request is pending

Two things capture immediately, whatever your organisation's approval state:

  • Repositories under the developer's own GitHub handle — a personal namespace is claimed by its owner, so there is no second gate to wait on.
  • Repositories with no GitHub remote at all — these are keyed locally to the machine rather than to a namespace, so no organisation owns them.

So a developer waiting on your approval is not locked out of the product. They are held at the boundary of one namespace, and everything outside it keeps working.

What you control and see

  • capture_policy for the organisation — approval_required, the default, or open.
  • The queue of pending requests. Each row identifies the requester by account id — a UUID, not a name or an email address — with your own row marked as yours. Resolving those ids to people is a later change.

That is the whole of it in this release. There is no per-repository override.

What you can retrieve

Everything captured is readable back by you, through the same record the audit surfaces read:

  • /app/commits indexes captured sessions by commit SHA, /app/conversations holds the sessions themselves, and /app/search runs full-text search across them.
  • /app/audit renders the record as an append-only log — entries in order, each linked to the one before it by a hash. The chain is tamper-evident: an entry edited after the fact no longer matches the link that points at it.
  • /app/audit/export produces a JSON bundle carrying the entries and every prev_hash / hash pair, so chain continuity can be recomputed from the bundle alone with no access to VibeCommit. The bundle states its own boundary: recomputing each hash from the original content needs the canonical record bytes, which it does not carry.
  • Your agent can query the same record while it works, through the MCP read tools.

What the record is, and what it is not

It is auditable, queryable and append-only: entries are added rather than rewritten, and the hash chain makes a later edit detectable. It is traceable in the sense of linkage— it binds a commit to the session recorded alongside it. That is a link between two records. It is not an account of what happened, and it is not a claim about what the agent did or did not do: the capture is the agent's own transcript of its own work, and the record's job is to keep that transcript in order and make changes to it afterwards detectable.

the code that runs on your machine

The plugin your coding agent loads is published at github.com/Vibe-Commit/claude-plugin. It vendors the exact binary that runs, so the build going onto your machine is the build you can inspect.