PixelRAG
A 100%-open-source visual retrieval system: instead of scraping a page into text and embedding chunks, it screenshots the page and retrieves the image, then a vision-language model reads the answer straight off the pixels. The bet: HTML parsing is where web RAG quietly loses information. Source: X/@akshay_pachaar, 2026-06-20
What it is
Traditional web RAG: fetch HTML → parse to text → chunk → embed → vector-retrieve → feed text to the LLM. Every step, especially parse-to-text, drops layout, tables, charts, and visual context. PixelRAG skips parsing entirely: render the page or document, capture screenshot tiles, index/retrieve over the image, and let a VLM answer from the rendered pixels. The official repo describes it as "Web Screenshots Beat Text" and ships code for webpages, PDFs, images, and a pre-built Wikipedia index. Source: StarTrail-org/PixelRAG README, 2026-06-30
Current source snapshot as of 2026-08-11: StarTrail-org/PixelRAG, Apache-2.0, Python, 9,416 stars, 807 forks, release/package v0.4.0, default-branch head c1dae49ba78e33ff1d02d9f8efd7b9adf9cecb22, and demo/API at pixelrag.ai. Version 0.4 adds bounded network-idle waiting, manifest-backed incremental reruns with atomic writes, stable article IDs, local image and Markdown/text rendering, CDP attachment, department filtering, device selection, traversal hardening, isolated worker profiles, and loop prevention. Source: GitHub repository/commit/release APIs; README; pyproject.toml; focused source tests, captured 2026-08-11
Install:
uv tool install pixelrag
The README exposes two practical surfaces:
pixelshot https://en.wikipedia.org/wiki/Python --output ./tiles
curl -X POST https://api.pixelrag.ai/search \
-H "Content-Type: application/json" \
-d '{"queries": [{"text": "What is the capital of France?"}], "n_docs": 5}'
Why it matters
- No parsing loss: tables, charts, multi-column layout, and styled content survive because the model sees the page as rendered, not as lossy extracted text.
- Less parser fragility: the indexed evidence no longer depends on one HTML-to-text extractor, but capture actions, authentication, render timing, and content moderation still need explicit handling.
- Rendered-state truth: JS-heavy components show up if the browser sees them. That is the real shift from "source text" to "screen state."
- A different point on the Retrieval Quality spectrum — image-native retrieval vs text-chunk retrieval. Complements, not replaces, text RAG (RAG System Architecture): use PixelRAG where the visual layout carries the answer.
Where it fits (browser/agent reading)
For agents that read the web, PixelRAG pairs with the rendering layer: a real browser (Browser Testing Skills) produces the screenshot; PixelRAG does the visual retrieval + VLM answer. The repo also ships a Claude Code plugin route named pixelbrowse: install pixelshot via uv tool install pixelrag or pipx install pixelrag, then install the marketplace plugin so Claude can call /screenshot https://example.com. Source: StarTrail-org/PixelRAG README, 2026-06-30
Use it when agents must extract from chart-heavy, table-heavy, dashboard, diagram, PDF, or JS-rendered pages where text extraction fails. The hosted API describes a pre-built index of 8.28M Wikipedia pages and accepts text or image queries; local pipelines can capture, chunk, embed, build a FAISS index, and serve search. The paper's 30M figure is screenshot tiles from roughly 7.13M content articles, not 30M pages. Source: StarTrail-org/PixelRAG README; PixelRAG paper Appendix A.2, captured 2026-08-11
Evidence and evaluation
The paper's strongest result is controlled, not universal. With Qwen3.5-4B and k=3, its Table 1 reports pixel-base accuracy above Trafilatura on six evaluated QA benchmarks. The largest base-model absolute gain is EVQA: 40.7 versus 29.6 (+11.1 points); the fine-tuned pixel model reaches 45.1 (+15.5 points). Natural Questions and NQ-Tables use a GPT-4.1 semantic judge in the published table, and strict exact match is materially lower. The abstract's "up to 18.1%" wording should therefore remain a paper claim, not a default production expectation. Source: PixelRAG paper §5.1–5.2; eval/README.md, captured 2026-08-11
The repository now includes a frozen evaluation client and endpoint preflights. Full reproduction still requires an H100 reader, several GPU retrieval services, roughly 570 GB of indexes, up to terabytes of tiles, external datasets, and an OpenAI key for judged cells. Kevin-Wiki did not reproduce Table 1. A focused local source test ran seven modules covering incremental reruns, network-idle behavior, stable article IDs, local text/image rendering, department filtering, and eval model configuration: 34 passed, 1 skipped. Source: PixelRAG eval/README.md, pyproject.toml, and local test receipt, 2026-08-11
Workflow recommendation
Route PixelRAG as one lane in a hybrid evidence contract:
- Retrieve exact/lexical and semantic text plus DOM/accessibility evidence when available.
- Add PixelRAG when layout, tables, charts, diagrams, or rendered state can change answerability.
- Preserve source revision, URL, tile coordinates, viewport, locale, authentication scope, waits, actions, and render timestamp as the visual hit identity.
- Preserve hyperlink targets and anchor text as structured metadata because pixels make links visible but not actionable.
- Compare per-lane recall, answer accuracy, citation/region grounding, latency, tokens, storage, privacy, and rights risk before choosing or fusing evidence.
Do not install it globally merely because a saved post is popular. Prefer an isolated uv tool/project trial, then promote only if a representative benchmark beats the existing text/browser route under the same reader, query set, and budget.
Failure boundaries
The bookmark thread is useful because it names the edge cases instead of overselling the method:
- Whatever state the browser captures is what gets indexed. Login-gated, personalized, collapsed, or hidden content still needs auth and scripted browser actions.
- Render timing matters. The renderer can wait for network idle or selectors, but paywalls and captchas remain hard.
- Visual retrieval avoids OCR/markdown round-trips during indexing, but the VLM still only reads retrieved tiles at answer time. Cost/latency can exceed cheap text embeddings.
- Site-specific actions can expand menus or click tabs before screenshotting, but that becomes per-site browser automation, not generic magic.
- Pixel representations lose actionable hyperlink structure unless link/anchor metadata is stored beside each tile.
- The paper reports nearly 6 TB of stored Wikipedia screenshots; render-on-demand trades that storage for runtime latency and reproducibility risk.
- Current evaluations are English-only, the Wikipedia-tuned retriever transfers imperfectly to news, and smaller/older VLM readers can underperform text retrieval.
- Pixel evidence can preserve harmful, misleading, copyrighted, or private content while evading string-based moderation. Rights, privacy, and content-policy checks belong before indexing and before agent delivery.
Timeline
- 2026-08-11 | Replayed the complete post, self-thread, video, current repository/release, paper, and evaluation package. Refreshed to
v0.4.0/c1dae49b, corrected the 8.28M-pages versus 30M-tiles distinction, recorded 34 passing focused tests, and promoted PixelRAG as a hybrid visual-retrieval lane with render-state, hyperlink, rights, moderation, storage, model-capability, and reproducibility gates. Source: X bookmark2068317780064276917capture manifest; StarTrail-org/PixelRAG current source and paper, 2026-08-11 - 2026-06-30 | Deep-reviewed the bookmark media, self-thread, repo, PyPI package, paper, and live API. Promoted PixelRAG from hub-only note into a versioned tool page:
v0.3.0, PyPI0.3.0, Apache-2.0, 5,702 stars, hosted API over 8.28M Wikipedia pages, andpixelshot/pixelbrowseagent-reading surfaces. Source: X/@akshay_pachaar, 2026-06-20; Source: StarTrail-org/PixelRAG, 2026-06-30 - 2026-06-22 | Captured from Kevin's X bookmark (@akshay_pachaar) as a new browser/retrieval tool — visual RAG that screenshots pages and reads answers off pixels via a VLM, skipping HTML parsing. Routed in the retrieval/browser stack. Source: X/@akshay_pachaar, 2026-06-20