LiteParse
LlamaIndex's open-source, model-free document parser: extract local text, Markdown, screenshots, structured data, and bounding boxes from PDFs, Office files, and images, with Rust, Node, Python, CLI, and browser/WASM surfaces. Source:
run-llama/liteparseHEADf7b7fb3, reviewed 2026-08-11
LiteParse is the “fast and light” parser from the team behind LlamaParse. It auto-detects a document's format, picks the parsing path, and returns local output with spatial layout. “World's fastest” and “most accurate” are vendor benchmark claims, not timeless routing facts; the captured June comparison and its limits are separated below. Source: X/@jerryjliu0, 2026-05-27 and 2026-06-18
What v2.0 changed
v1 shipped only as a Node/TypeScript package, which capped latency and distribution. v2.0 rewrote the project in Rust with a single core that propagates to every language binding, so LiteParse now runs as a native Rust crate, Node/TS package, Python package, and a WASM build that runs in the browser and on edge runtimes. Source: LlamaIndex blog "Up to 100x Fast Parsing with LiteParse v2.0 and Rust", https://www.llamaindex.ai/blog/liteparse-v2-0-runs-everywhere, 2026-06-12
| Surface | Install |
|---|---|
| Node / TypeScript | npm i @llamaindex/liteparse (or -g for the CLI) |
| Python | pip install liteparse |
| Rust (lib + CLI) | cargo add liteparse / cargo install liteparse |
| Browser / edge (WASM) | npm i @llamaindex/liteparse-wasm |
How it works
- Hybrid extraction: pull structured embedded text directly from the file, fall back to OCR for scanned regions, and route OCR through bundled Tesseract or an HTTP OCR server such as EasyOCR or PaddleOCR. Source: run-llama/liteparse README, 2026-06-30
- Engine: a custom fork/build of PDFium for PDF rendering plus
tesseract-rsas the default OCR backend; Rust ownership of the build pipeline is what unlocks the WASM/edge target. Source: LlamaIndex blog, 2026-06-12 - Formats: PDF; Microsoft Office (
.docx/.xlsx/.pptx) and OpenDocument (.odt/.ods/.odp) via LibreOffice; images (.png/.jpg/.tiff) via the OCR path. Source: crates/liteparse README, 2026-06-12 - Output: text, Markdown, screenshots, and JSON with bounding boxes, suitable for downstream chunking/RAG or visual page checks. Source: run-llama/liteparse README, 2026-06-30
Markdown and benchmark evidence
The complete v2 launch image is narrower than the headline. After five warmup
runs with OCR disabled, LiteParse v2 is fastest on the shown 1-, 60-, and
457-page fixtures (0.002 s, 0.123 s, and 0.777 s), but PyMuPDF wins the
24-page fixture (0.029 s versus 0.041 s). Hardware, exact dependency
revisions, variance, and output-parity checks are absent, and the compared tools
do not all emit identical semantics. The image supports a large launch-era Rust
overhead reduction on that fixture set—not “fastest on every document.”
Source: X/@jerryjliu0 v2 launch image and LlamaIndex v2 blog, replayed
2026-08-11
Version 2.1 added heuristic PDF-to-Markdown classification using PDF font,
position, and grid-projection signals to identify paragraphs, headings, lists,
tables, images, and links. The complete local 26.5-second browser demo parses a
NASA PDF and switches the same result among text, Markdown, and JSON with
coordinates. It demonstrates WASM-local interaction and the three output modes;
it does not prove comparative accuracy, speed, accessibility, memory use, or
browser coverage. Source: LlamaIndex “Markdown Comes to LiteParse”; local
video wiki/assets/x-bookmarks/2067679507126124858/video-01.mp4, reviewed
2026-08-11
Fresh enrichment recovered all four self-thread benchmark screenshots:
| Captured release comparison | LiteParse result | Important boundary |
|---|---|---|
| ParseBench | 0.328 overall |
competitors lead some semantic-formatting, chart, and visual-grounding columns; the source calls model-free chart/grounding results effectively noise |
| olmOCR-bench | 39.2% overall |
strong structural categories, but OpenDataLoader leads long-tiny-text and all compared tools are at or near zero on math/old-scan math |
| OpenDataLoader-bench | 0.871 overall |
0.908 reading order, 0.693 tables, 0.816 headings in this historical five-parser run |
| Fixed-set speed | 3.16 ms/page |
the screenshot omits hardware, warmup, confidence intervals, and exact dependency revisions |
This supports “led this disclosed model-free comparison at release,” not an
eternal category winner. The current OpenDataLoader benchmark README at
7af1d8f no longer includes LiteParse in its result table, so current adoption
must run representative documents against current candidates. The source and
benchmark images are preserved in the capture manifest. Source: X/@jerryjliu0
self-thread, 2026-06-18; current OpenDataLoader benchmark, reviewed 2026-08-11
Current authority as of 2026-08-11: LiteParse repository HEAD and WASM tag
2.11.1 are f7b7fb3b6de4981cbad91b64df4af09384b05a6c; npm
@llamaindex/liteparse, npm @llamaindex/liteparse-wasm, PyPI liteparse,
and the Rust crate manifest all report 2.11.1. The repository is Apache-2.0.
Source: GitHub, npm, PyPI, and captured Cargo manifest, 2026-08-11
Agent skill route
LiteParse has an upstream Agent Skill in run-llama/llamaparse-agent-skills: npx skills add run-llama/llamaparse-agent-skills --skill liteparse. The main operational point is not “parse everything”; it is parse once, then search bounded windows rather than repeatedly dumping a full document into context. Kevin's executable owner adds an output-mode route: text for search, Markdown for document structure, and JSON for coordinates/provenance. Current upstream HEAD 2dcef7c has no LiteParse-file diff from tag liteparse-1.0.1, so no blind skill refresh is needed. Source: captured upstream skill repository, reviewed 2026-08-11
LiteParse vs LlamaParse vs the parser field
LiteParse deliberately stops at fast, local, model-free extraction. For dense tables, multi-column layouts, charts, handwriting, or low-quality scans, LlamaIndex points users at the cloud LlamaParse product, which adds model-driven structure for production pipelines. Source: run-llama/liteparse README, 2026-06-12
Within Kevin's parsing shelf:
- OmniParse — heavier, GPU-friendly multimodal ingestion (Surya OCR + Florence-2 + Whisper, web crawling). Use when the corpus is image/audio/video-heavy.
- MarkItDown - Universal File-to-Markdown Converter — Microsoft's lightweight pure-Python doc→markdown converter; closest in spirit but Python-only and less spatially precise.
- LiteParse — the speed/portability pick: one Rust core, runs in CI, agent workflows, the browser, or edge, with no tokens spent. Feeds RAG pipelines the same way.
Use LiteParse first for local PDF/layout extraction when speed, privacy, or screenshots matter. Use MarkItDown - Universal File-to-Markdown Converter when broad ordinary file-to-Markdown conversion is enough; escalate to MinerU for hard OCR/table/formula PDFs and OmniParse for multimodal/GPU-heavy corpora. Run the three-gate chain (Brin -> skill-auditor -> reputation) before adopting new external packages in production, per AGENTS.md.
Timeline
- 2026-08-11 | Force-replayed the earlier v2 launch source and complete
benchmark image. Added the PyMuPDF-winning counterexample, five-warmup/OCR-off
conditions, output-parity boundary, and OCR/LlamaParse escalation replies;
current
2.11.1source remains unchanged. Source: X/@jerryjliu0 v2 launch; LlamaIndex v2 blog; current repository - 2026-08-11 | Recovered and inspected all four benchmark screenshots plus the complete browser demo; reframed the headline as a historical disclosed comparison, added explicit competitor and hard-document limits, refreshed all package surfaces to
2.11.1, and registered LiteParse as a stable capability and executable skill object. Source: X/@jerryjliu0 self-thread; GitHub/npm/PyPI; local artifacts, reviewed 2026-08-11 - 2026-05-27 | @jerryjliu0 (LlamaIndex) announced LiteParse v2 — full Rust rewrite, "world's fastest PDF parser," more accurate than model-free parsers, completely open-source and free. 4,076 likes, 4,561 bookmarks. Source: X/@jerryjliu0, 2026-05-27
- 2026-06-12 | Dedicated page created from dev-tools bookmark absorption; researched repo (
run-llama/liteparse, Apache-2.0, ~9K stars) + LlamaIndex v2.0 blog. Source: x-bookmark-absorb, 2026-06-12 - 2026-06-30 | One-by-one bookmark artifact review inspected the launch benchmark image, self-thread, OCR reply, upstream repo, npm/PyPI/Git tags, and the upstream Agent Skill. Refreshed routing to make LiteParse the fast local PDF/layout route and imported the skill into
skills/engineering/liteparse/. Source: X bookmark artifact audit, 2026-06-30 - 2026-07-01 | Refreshed package/source snapshot: npm
@llamaindex/liteparseand@llamaindex/liteparse-wasmadvanced to2.4.0; GitHub tagdocker-v2.4.0resolves to current HEADa8288d09cb6bf93c2f7b5257e9634afd26a4566b. Source: npm registry packument;git ls-remote, 2026-07-01 - 2026-06-30 | Added the later benchmark/demo bookmark: self-thread claims accuracy and speed tests; local video shows the browser parser producing text, Markdown, and JSON from a PDF. Version snapshot was
2.3.0. Source: X/@jerryjliu0, 2026-06-18; Source: npm/PyPI/GitHub, 2026-06-30