Source Compile Workflow

Sub-daily proposal loop that syncs raw capture sources, extracts and interrogates signals, and delivers evidence-bound owner changes to Daily Brief. A separate exact-approval path performs canonical writeback.

The source-compile automation turns accumulated raw data into reviewable proposals every four hours. It is the scheduled entry point to the shared Wiki Brain Operating Model spine, not a bookmark-only digest. The research run owns capture -> expand -> correlate -> extract -> bind -> interrogate -> propose; exact approval and a separate apply path own write back -> prove. Source: automations/source-compile.md, 2026-05-31; User request, 2026-06-30; User request, 2026-08-10; Automation V2 audit, 2026-08-11

Trigger And Inputs

Run every four hours through the local scheduler, manually after a large capture batch, or when scripts/check-freshness.ts flags stale X bookmarks / source compile. Required inputs:

  • scripts/sync-x-bookmarks.sh
  • scripts/enrich-x-bookmarks.ts
  • scripts/download-x-bookmark-media.ts
  • scripts/export-ai-sessions.sh
  • skills/productivity/x-bookmark-absorb/SKILL.md
  • skills/productivity/absorb-sources/SKILL.md
  • brain/{goals,loops,objects,sources}.json
  • reviews/source-signals/*.json
  • state.json compile markers
  • raw bookmark/enriched records under raw/x-bookmarks/

Local Codex is the preferred runner because it can access local browser-authenticated sources and local session transcripts. Cloud runners can only process committed raw/enriched records.

Phases

  1. Sync and index : Run ./scripts/sync-x-bookmarks.sh, npx tsx scripts/enrich-x-bookmarks.ts, npm run download:x-bookmark-media -- --synced-date "$(date +%F)", and npm run brain -- sessions index. Bookmark artifacts may be frozen locally; Codex, Cursor, and Claude evidence files are indexed in their provider stores and are not copied into raw/. The private session index records stable identity, native session ID, provider, workspace hint, first user-prompt preview, byte count, modification time, fingerprint, trace kind, and parent lineage under .brain/sessions/. It reports top-level conversations separately from delegated subagent/workflow traces, so a provider's internal fan-out never inflates the operator's task count. [Sources: automations/source-compile.md; packages/brain-runtime/src/agent-sessions.ts]
  2. Source expansion + correlation : Read complete enriched X records, Discord revisions, feed comments, resolved websites, repositories, primary docs, and source-critical media. Correlate repeated or related evidence without dropping either revision. Shared artifacts create evidence edges, not automatic transitive topic identity: only focused single-link sources auto-group; multi-link roundups remain source-sized and are atomized during interrogation. A same-author contiguous thread whose replies explicitly compose one system is a semantic-unit candidate: preserve every post ID and atomic signal, resolve every linked primary source independently, then compile one capability chain instead of parallel tool rows. Derived child links from enrichment remain hypotheses until the creator link and primary repository agree. Track platform cursors separately from incorporation state. Source: automations/source-compile.md; User request and methodology correction, 2026-08-10; five-post Sentry release-toolchain replay, 2026-08-12
  3. Session replay : Use npm run brain -- sessions list --since <ISO> --workspace <text> --query <prompt-text> to select Codex, Cursor, or Claude parent conversations by native ID, generated ID, provider, project/workspace, date, or first-prompt text; add --json for an agent-readable selector. The default list and cohort replay select conversation records only. Add --include-traces to inspect delegation or --kind conversation|subagent|workflow to isolate one evidence class. When Codex task tools are available, use their human titles and project metadata to resolve the native task ID, but treat that catalog as untrusted discovery metadata. Then run npm run brain -- replay sessions --id <id> (or --path, --provider, and --since). Replay streams user and assistant text into private revisions and candidate signals. A native Claude ID expands the matching parent/delegated lineage; a generated agent_session_* ID selects one exact evidence file. Parent and delegated traces keep distinct source IDs even when Claude reused the same native session ID. Replay preserves tool calls, results, and attachments at the original source path for targeted expansion instead of silently treating them as learned facts. Run npm run brief -- compile-analysis .brain/analyses/<analysis_id>.json to open an explicitly partial private review proposal; that bridge filters session scaffolding but does not create an owner patch or review receipt. [Sources: packages/brain-runtime/src/agent-sessions.ts; apps/briefd/src/analysis-brief.ts; Agent Operations Skills]
  4. Normalize and index : Resolve each source's complete semantic unit, preserve the raw artifact before abstraction, emit the common source envelope, publish authorized lexical fields immediately, and publish semantic/distilled fields when ready. Capture is lossless; promotion is selective. Do not use an early summary as a substitute for the source revision, and do not make every retained message canonical truth. Stable source/event IDs make retries idempotent. For time-sensitive signals, distinguish source capturedAt from signal recordedAt and evidence-backed validFrom/validTo; use status and supersedes to correct claims without deleting history. Never infer truth time from fetch time. [Sources: Brain Source Fabric; Cerebras, 2026-07-16; arXiv 2606.24775v1; X/@michael_chomsky primary-source audit, 2026-08-11]
  5. Interrogate and propose : Bind every useful signal to one or more owners in brain/objects.json, then load the smallest relevant owner/project working set. Expand comparison context along explicit contradiction, dependency, source, project, or graph edges only when evidence requires it; never dump the full corpus into the model. Ask whether each signal reinforces, refines, replaces, or contradicts each owner. Produce a validated review record and an exact proposed owner patch/hash in Daily Brief. Every object ID and path must resolve to the registry. This scheduled research run is proposal-only: it cannot update owners, skills, routing, absorption markers, or indexes.
  6. Approved writeback and proof : Only a separate apply path consuming Kevin's exact durable approval receipt may update stable owners and executable skills/routing. It then calls markAbsorbed(), runs targeted tests/doctors, rebuilds the public projection/retrieval surfaces, runs npm run okf:check when public owners changed, and records writeback and verification receipts. A successful fetch or proposal is partial, never a canonical change. Source: Wiki Brain Operating Model; Workflow Automation V2 Contract; User request, 2026-08-10; OKF v0.2 SPEC

Two-Speed Review

Admission and adoption are separate proof obligations. A complete source does not need to wait for every repository, benchmark, vendor claim, or current version to be deeply researched before its actual signal can enter the evidence graph.

  1. Admission batch: read the complete available source unit, preserve its revision/hash, extract every useful atomic signal, bind each signal to stable objects, interrogate those owners, and record reinforce or watch when no writeback is justified. Text-only sources may be reviewed in coherent batches of at most 50; every counted source must appear in at least one signal. For message corpora, every available message first receives a stable source identity. A lineage may then close with one exact review-unit receipt whose digest binds every current message source, revision, and signal; the compact review may cite representative safe evidence without publishing every body. This is aggregation after identity and classification, not sampling.
  2. Deep adoption: schedule primary-source, repository, security, benchmark, license, current-version, or visual-artifact work whenever a recommendation, installation, factual promotion, or behavior change depends on it. Admission never grants default status.
  3. Durable obligation: compile review followUps into .brain/queues/deep-source-followups.json and brain/deep-review-queue.json. Regenerating the admission queue cannot erase these tasks after their sources become integrated. Complete adoption append-only: a new proof-bearing review names the exact queue IDs in completeTaskIds; the compiler resolves each immutable source, reason, and target tuple into a complete follow-up. Never edit the admission review to manufacture completion.
  4. Artifact lanes: source-critical images and video require contact-sheet or direct visual inspection; incomplete enrichment enters recovery. Neither may be counted through the text lane.

Run npm run review:batches to plan bounded lanes, compile an authored decision with npm run review:batch:compile -- <decision.json>, then run npm run review:deep-queue. The fast path changes scheduling, never the evidence contract. Receipt-backed source ordinals keep large batches exact; native-source materialization applies the same review contract to X, Discord, and Hacker News; recovered X artifacts re-enter the visual lane instead of remaining false recovery work. The current admission queue is complete at 1,945/1,945 integrated units. The durable queue currently contains 240 tasks over 774 retained sources: 235 complete and 5 still queued after the font and release-thread completions. Source: brain/deep-review-queue.json, generated 2026-08-12; User throughput corrections, 2026-08-11 and 2026-08-12; batch-decision receipt, native-source, recovery-artifact, and queue contract tests

Deep adoption should also run as destination-coherent cohorts, not as one agent turn per bookmark. Group 12–24 queued tasks, normally no more than 100 sources, when they interrogate the same small owner set and can share primary documentation, repository receipts, or a frozen evaluation protocol. Research the shared evidence once, preserve source-level coverage and limitations, apply one consolidated owner diff, compile the decision, and validate the entire cohort. Split a task only when its authority, visual artifact, benchmark, or security boundary is genuinely independent. Throughput is measured as tasks closed, sources resolved, owner changes proved, and queue remaining—not pages written or links opened. Source: User speed correction, 2026-08-12; OpenAI, “Harness engineering: leveraging Codex in an agent-first world,” 2026-02-11

For an authored system thread, destination coherence also applies inside the thread. Build a stage graph before deciding: name each linked component's job, inputs/outputs, conditions, authority, and proof; identify the incumbent and omittable stages; then update the router/workflow once. Engagement belongs to the whole reconstructed signal, not to each reply as a competing ranking row. The Sentry release thread is the reference: Stricli, a generated API client, Craft, Fossilize, and binpatch are conditional CLI/API/release/package/update stages, not a five-tool default. Source: X 2082947912498024751– 2082947922648281278; exact-revision repository explorations

What to look for

The compiler is not a news digest. It is a promotion system. Promote only signals that change future behavior or durable knowledge:

  • New tools Kevin is likely to use.
  • Repos, docs, demos, papers, templates, datasets, or artifacts posted by the author in replies.
  • API/framework changes that invalidate a skill, workflow, or style page.
  • Repeated agent mistakes that should become guardrails.
  • Project facts that change the current status, architecture, metrics, or positioning.
  • People, companies, papers, or concepts that recur across sources.
  • Contradictions between old wiki claims and newer evidence.

Every saved source can carry preference, context, or routing signal. If it cannot change an owner today, keep it held with its identity, rationale, and reopening condition. The goal is compounding without inventing a new page for every observation.

Compilation rules

  • Search before creating pages.
  • Use enriched bookmark records, not raw tweet text alone.
  • Inspect artifactReplies before deciding a bookmark is linkless or low-signal.
  • Research promoted tools/projects beyond X when primary or independent sources exist.
  • Prefer updating existing compiled truth over adding a disconnected note.
  • Update every materially affected stable owner, but do not manufacture page count.
  • Add timeline entries only for dated evidence.
  • Use markAbsorbed() for raw files so wiki-status does not keep surfacing them.
  • If the source changes how agents should operate, update skills/, wiki/SKILL-RESOLVER.md, config/cursor/rules/, or wiki/meta/.
  • If the source changes project status, update the project page and any linked career/profile page.
  • Treat events as refresh triggers, not necessarily retrievable documents: refetch a complete thread, article artifact, repository symbol/file, or other semantic unit before upsert.
  • Keep lexical indexing, semantic distillation, canonical compilation, and human review as separate states so one failure does not falsely block or complete the others.
  • Apply source ACL and workflow/project scope before retrieval, fusion, reranking, or context expansion.
  • Never union review units through a multi-link roundup. Preserve its links as separate evidence edges, extract every useful item, and let explicit object comparisons—not graph transitivity—determine joint review.
  • Reconstruct same-author system threads before queuing deep work. Preserve every post as source evidence, but close duplicate stage-level obligations with one later proof-bearing completion review and one consolidated owner diff. Do not edit the original admission review to manufacture completion.
  • Verify enriched child-link identity against the creator's actual URL and the primary repository. A guessed, moved, or nonexistent child repository is not evidence and must not enter routing.
  • Treat media as artifact proof rather than a topic key. Two posts do not become the same subject merely because their images or videos share delivery infrastructure.
  • Treat brain/objects.json as the foreign-key registry for reviewed operating owners. brain doctor must fail when a review names an absent object ID or when the review path differs from the registered path.
  • Preserve source context at write time and filter late. The archive may retain every message/revision while stable owners receive only interrogated, source-bound changes.
  • Never content-deduplicate messages. Retry one stable delivery idempotently, correlate repeated text, and retain every distinct message source ID. A transcript snapshot is custody evidence, not a replacement for message-level identity.
  • Bind aggregated review closure to the exact source/revision/signal digest. integrated and held are final for that digest; partial and blocked remain pending, and any changed evidence automatically reopens the unit.
  • Treat early localization and complete evidence assembly as separate retrieval targets; a plausible top result does not complete a multi-source review.
  • Prefer local, digest-bound owner maintenance over broad canonical rewrites. Rebuild derived indexes when needed, but do not confuse reindexing with semantic integration.
  • Preserve conflict, missing-evidence, and no-answer states. A source compiler is allowed to hold or abstain; it is not allowed to invent a stable conclusion so the queue looks complete.
  • Do not turn deep research into an admission gate. Preserve and bind complete source text now; schedule the narrower primary/repository/security/current- state work that controls adoption in the durable deep queue.
  • Give time-sensitive signals an optional validated temporal envelope: recordedAt, validFrom, validTo, status, and supersedes. Source capture time is provenance, not a claim-validity interval.
  • Treat vendor build-versus-buy claims as incentive-bearing evidence. Verify cited examples at their primary sources, distinguish portable differentiation from commodity operations, and require an export/replay path before a hosted execution provider can become a workflow default.
  • Treat generated wikis as derived projections. OpenWiki's code/personal modes, source connectors, owned Markdown, managed pointer blocks, update/no-op flow, graph viewer, and OKF output are useful reference behavior, but generated docs do not inherit human intent or canonical authority. Pin the source revision, preserve project-owned instructions, declare telemetry/provider/credential paths, diff updates, cite source files/traces, and prove deletion/drift before adopting it for one repository. Source: langchain-ai/openwiki@c5d41cb, reviewed 2026-08-12
  • Prefer bounded, project-specific navigation commands when they answer a repeated repository question more precisely than general indexing. A discover.sh-style wrapper may combine rg, Git, glossary/spec/project rules, outlines, callers, and a transparent incremental JSONL index; document its false positives, freshness and rebuild rule, then compare it with raw rg, qmd, and Graphify before promotion. Source: Seb/plainionist, “discover.sh,” 2026-07-11
  • Rich media capture is a source connector, not memory truth. Bloom preserves a local recording/library while streaming capture to VideoDB for export, transcription, indexing, hosted chat, and share links. Record screen/mic/ system-audio consent, API/provider egress, collection/video IDs, local and remote deletion, share revocation, silent-video behavior, transcript/timestamp fidelity, recovery, cost, and rights before compiling any recording. Never repeat “local-first” as “local processing.” Source: video-db/bloom@65aa9bf; VideoDB build note, reviewed 2026-08-12
  • Treat memory, generated-wiki, vault-rewrite, and structured-extraction tools as noncanonical projections until they prove source ownership, immutable revisions, citations or locators, ACL enforcement before retrieval, conflict handling, deterministic replay, correction/deletion, privacy and provider egress, tests, export, and digest-bound human-approved writeback. Claude-Mem, GBrain, Obsidian rewrite agents, AutoWiki, LangExtract, semantic extractors, and Markdown editors are retained because each carries useful patterns; none may replace the raw evidence store or stable owner by convenience. A current repository receipt proves inspectable code, not those end-to-end properties. Source: X 2033207845521707471, 2043234633408798932, 2044291663213015491, 2061225218547269909, 2067325077700616459, 2067749509313302866, 2039726865775051135, 2067749509313302866, 2074020481729286181, 2076991032487776491; current Claude-Mem, Obsidian Second Brain, and LangExtract repository receipts, reviewed 2026-08-12
  • A missing attachment, parent conversation, article body, or repository is a preserved evidence-gap state, not permission to reconstruct the claim from a teaser. Keep the source ID, available text/media metadata, attempted routes, target owners, and reopening condition; make no substantive promotion until the semantic unit is recovered. Source: X 2029045945980305732, 2068497127140213151; source recovery review, 2026-08-12
  • Treat a social-video caption, watermark, and timeline as claims about an artifact, not as the artifact's identity. For a clipped or recut talk, match frames, spoken phrases, duration/offset, speaker order, and event metadata to the publisher's original recording; then split every age, employer, compensation, revenue, outcome, and causal statement into a separately sourced claim. A university logo proves where footage was recorded, not the overlay's biography or business story. If the original transcript does not contain the claim, preserve the recut as a discovery lead and misinformation/copywriting counterexample while compiling the actual talk on its own merits. Repeated invented overlays from an account are a source-risk prior, not permission to skip inspection: keep useful linked artifacts, but require primary evidence for each consequential claim. Source: X/@noisyb0y1 2087128585144357274; Stanford CS25 V4 schedule; Stanford CS25 original recording, reviewed 2026-08-12
  • Treat file classification, conversion, transcription, OCR, and semantic extraction as separate derived revisions. Record original digest and type, classifier and converter/model versions, options, provider/network path, page/slide/sheet/speaker/time locators, preserved versus omitted structure and media, confidence/disagreement, failed elements, cost, update/deletion state, and the exact derived digest. A Markdown or transcript projection helps retrieval; it never replaces the original or proves parser safety and semantic fidelity. Source: current Magika, MarkItDown, pdf-brain, TranscriptionSuite, VibeVoice, and OfficeCLI sources; reviewed 2026-08-12
  • A fast multi-format converter is a candidate projection engine, not the source of truth. For mixed legacy Office inputs, AnyDoc is worth a pinned comparison because it recognizes content, exposes typed unsupported, malformed, encrypted, resource-limit, missing-part, and I/O failures, shares a document model across most non-PDF formats, and has Rust, Node, Python, WASM, and CLI surfaces. The adapter must still preserve the original bytes and digest; record detected versus declared type, converter revision and options, conversion errors, omitted media/structure, source locators, and output digest; and route scanned/image-only PDFs to OCR. AnyDoc's published speed/quality chart uses a private 100-document corpus, excludes PDF from the comparison, compares trigrams against AnyDoc rather than neutral truth, gives competitors different process-boundary treatment, and averages different supported-format sets. Keep its per-format evidence as a hypothesis until an owned corpus reproduces it. Source: firecrawl/anydoc@4e3089b1, README, parsers, errors, CI, tests, fuzz targets, skill, and benchmark harness; reviewed 2026-08-12

Promotion Matrix

Source compile signal Durable action
Tool/project/API with docs, repo, demo, paper, or repeated bookmark signal Update or create tool/project/concept page, then link from the relevant hub.
Bookmark has source-critical media or author artifact reply Preserve local media/artifact cards and cite the enriched record.
Repeated agent friction across 2+ sessions Update/create executable skill or workflow, then regenerate skill registry when needed.
Contextual or currently non-actionable source Keep a review record as held with candidate owners and a reopening condition; do not create a stub page.
Existing page updated today Avoid churn unless new evidence corrects or materially changes a claim.
Sync failure or incomplete enrichment Record blocker and do not advance compile markers past unprocessed records.

Outputs

  • Wiki page updates with [Source: ...] citations
  • New skills when a pattern repeats 2+ times
  • Append to wiki/log.md
  • npx tsx scripts/build-index.ts and qmd update && qmd embed after substantive edits
  • npm run okf:check after public canonical owners or privacy rules change
  • state.json freshness updates from the sync and index scripts
  • wiki/_absorb_log.json updates for raw-source absorption
  • state.json updates for compile.last_bookmark_id, compile.last_bookmark_synced_at, and ISO automation freshness
  • .brain/queues/deep-source-followups.json and brain/deep-review-queue.json when review follow-ups change

Output Report Shape

Every scheduled run should leave outputs/<YYYY-MM-DD>/source-compile/<agent>.md with:

  • sources synced and commands run
  • bookmark/session slice reviewed
  • durable promotions made
  • explicit deferrals and why they were deferred
  • owner bindings, interrogation outcomes, and review-record paths
  • raw files marked absorbed
  • generated surfaces refreshed
  • remaining backlog or next cursor

If the run finds no promotable source, say No durable wiki update needed and name the sources checked. A silent no-op is indistinguishable from a skipped run, which is exactly how stale brains learn to lie.

State Writes

source-compile has three state layers:

State key Update when
syncs.x_bookmarks X bookmark sync succeeds.
compile.last_bookmark_synced_at / compile.last_bookmark_id Bookmarks through that point have validated integrated, held, or blocked review records.
Source/repository assembled revision The known affected semantic units at that exact source revision were captured, interrogated, written back or explicitly held, and proved.
Source/repository reconciliation revision A separate omission audit re-read the complete current source surface and checked dependent owners; ordinary incremental sync must not advance it. Record this revision in the review/proof layer until a connector has a dedicated typed field.
automations.source-compile The scheduled automation completed its check, using an ISO timestamp.

Do not advance compile markers just because sync succeeded. The marker means the wiki has reviewed the record.

Likewise, do not call a source “fully reconciled” merely because path-aware or event-aware incremental work is current. Freshness answers whether known affected material was updated; reconciliation answers whether an omission audit looked for material the dependency map did not know to request. This two-clock rule is adapted from ralph-vault-skill's separate last_sync_commit and last_reconcile_commit; Kevin's authoritative evidence remains the source revision and validated review receipt. Source: SantanderAI/ralph-vault-skill@c5eff8c, reviewed 2026-08-11

Validation

After a source-compile run:

npx tsx scripts/build-index.ts
npx tsx scripts/doctor.ts --quiet
npm run okf:check
qmd update
qmd embed

Then inspect freshness:

npx tsx scripts/check-freshness.ts

If the same source still appears as pending, the absorb log was not updated or the source was only partially compiled.

Failure Handling

Failure Response
ft sync or X auth fails Record command/error in output; leave X sync and compile markers unchanged.
Enrichment fails for some records Process only records with enough source context; defer failed rows explicitly.
Duplicate event or retried delivery Acknowledge safely, deduplicate by stable event ID, and do not create a second evidence row or canonical effect.
Semantic distillation is delayed Keep authorized raw/normalized text lexically searchable; mark the semantic projection pending rather than failing the entire source.
Partial semantic-unit fetch Do not overwrite a complete prior unit with an incomplete snapshot; retry or dead-letter with cursor and source identity intact.
ACL or project-scope metadata is missing Quarantine before indexing; never infer broad visibility.
Media download fails Preserve source tweet/enriched record; mark media repair/deferral in the artifact audit when relevant.
qmd/build-index fails after wiki writes Report failure and leave exact command output; do not claim retrieval is current.
Too many new bookmarks for one run Process a coherent slice, record remaining count, and leave markers at the last processed/deferred record.

Manual trigger

Tell any agent: "Run the source-compile automation." Or: npx tsx scripts/run-automation.ts source-compile.

Run Contract

This workflow follows Workflow Run Contract: name the sources, write output or no-op proof, promote only durable facts, update state only when the run really completed, refresh generated surfaces when durable pages or skills change, and log user-visible work.


Timeline