LiteParse

LlamaIndex's open-source, model-free document parser: extract local text, Markdown, screenshots, structured data, and bounding boxes from PDFs, Office files, and images, with Rust, Node, Python, CLI, and browser/WASM surfaces. Source: run-llama/liteparse HEAD f7b7fb3, reviewed 2026-08-11

LiteParse is the “fast and light” parser from the team behind LlamaParse. It auto-detects a document's format, picks the parsing path, and returns local output with spatial layout. “World's fastest” and “most accurate” are vendor benchmark claims, not timeless routing facts; the captured June comparison and its limits are separated below. Source: X/@jerryjliu0, 2026-05-27 and 2026-06-18

What v2.0 changed

v1 shipped only as a Node/TypeScript package, which capped latency and distribution. v2.0 rewrote the project in Rust with a single core that propagates to every language binding, so LiteParse now runs as a native Rust crate, Node/TS package, Python package, and a WASM build that runs in the browser and on edge runtimes. Source: LlamaIndex blog "Up to 100x Fast Parsing with LiteParse v2.0 and Rust", https://www.llamaindex.ai/blog/liteparse-v2-0-runs-everywhere, 2026-06-12

Surface Install
Node / TypeScript npm i @llamaindex/liteparse (or -g for the CLI)
Python pip install liteparse
Rust (lib + CLI) cargo add liteparse / cargo install liteparse
Browser / edge (WASM) npm i @llamaindex/liteparse-wasm

How it works

Markdown and benchmark evidence

The complete v2 launch image is narrower than the headline. After five warmup runs with OCR disabled, LiteParse v2 is fastest on the shown 1-, 60-, and 457-page fixtures (0.002 s, 0.123 s, and 0.777 s), but PyMuPDF wins the 24-page fixture (0.029 s versus 0.041 s). Hardware, exact dependency revisions, variance, and output-parity checks are absent, and the compared tools do not all emit identical semantics. The image supports a large launch-era Rust overhead reduction on that fixture set—not “fastest on every document.” Source: X/@jerryjliu0 v2 launch image and LlamaIndex v2 blog, replayed 2026-08-11

Version 2.1 added heuristic PDF-to-Markdown classification using PDF font, position, and grid-projection signals to identify paragraphs, headings, lists, tables, images, and links. The complete local 26.5-second browser demo parses a NASA PDF and switches the same result among text, Markdown, and JSON with coordinates. It demonstrates WASM-local interaction and the three output modes; it does not prove comparative accuracy, speed, accessibility, memory use, or browser coverage. Source: LlamaIndex “Markdown Comes to LiteParse”; local video wiki/assets/x-bookmarks/2067679507126124858/video-01.mp4, reviewed 2026-08-11

Fresh enrichment recovered all four self-thread benchmark screenshots:

Captured release comparison LiteParse result Important boundary
ParseBench 0.328 overall competitors lead some semantic-formatting, chart, and visual-grounding columns; the source calls model-free chart/grounding results effectively noise
olmOCR-bench 39.2% overall strong structural categories, but OpenDataLoader leads long-tiny-text and all compared tools are at or near zero on math/old-scan math
OpenDataLoader-bench 0.871 overall 0.908 reading order, 0.693 tables, 0.816 headings in this historical five-parser run
Fixed-set speed 3.16 ms/page the screenshot omits hardware, warmup, confidence intervals, and exact dependency revisions

This supports “led this disclosed model-free comparison at release,” not an eternal category winner. The current OpenDataLoader benchmark README at 7af1d8f no longer includes LiteParse in its result table, so current adoption must run representative documents against current candidates. The source and benchmark images are preserved in the capture manifest. Source: X/@jerryjliu0 self-thread, 2026-06-18; current OpenDataLoader benchmark, reviewed 2026-08-11

Current authority as of 2026-08-11: LiteParse repository HEAD and WASM tag 2.11.1 are f7b7fb3b6de4981cbad91b64df4af09384b05a6c; npm @llamaindex/liteparse, npm @llamaindex/liteparse-wasm, PyPI liteparse, and the Rust crate manifest all report 2.11.1. The repository is Apache-2.0. Source: GitHub, npm, PyPI, and captured Cargo manifest, 2026-08-11

Agent skill route

LiteParse has an upstream Agent Skill in run-llama/llamaparse-agent-skills: npx skills add run-llama/llamaparse-agent-skills --skill liteparse. The main operational point is not “parse everything”; it is parse once, then search bounded windows rather than repeatedly dumping a full document into context. Kevin's executable owner adds an output-mode route: text for search, Markdown for document structure, and JSON for coordinates/provenance. Current upstream HEAD 2dcef7c has no LiteParse-file diff from tag liteparse-1.0.1, so no blind skill refresh is needed. Source: captured upstream skill repository, reviewed 2026-08-11

LiteParse vs LlamaParse vs the parser field

LiteParse deliberately stops at fast, local, model-free extraction. For dense tables, multi-column layouts, charts, handwriting, or low-quality scans, LlamaIndex points users at the cloud LlamaParse product, which adds model-driven structure for production pipelines. Source: run-llama/liteparse README, 2026-06-12

Within Kevin's parsing shelf:

  • OmniParse — heavier, GPU-friendly multimodal ingestion (Surya OCR + Florence-2 + Whisper, web crawling). Use when the corpus is image/audio/video-heavy.
  • MarkItDown - Universal File-to-Markdown Converter — Microsoft's lightweight pure-Python doc→markdown converter; closest in spirit but Python-only and less spatially precise.
  • LiteParse — the speed/portability pick: one Rust core, runs in CI, agent workflows, the browser, or edge, with no tokens spent. Feeds RAG pipelines the same way.

Use LiteParse first for local PDF/layout extraction when speed, privacy, or screenshots matter. Use MarkItDown - Universal File-to-Markdown Converter when broad ordinary file-to-Markdown conversion is enough; escalate to MinerU for hard OCR/table/formula PDFs and OmniParse for multimodal/GPU-heavy corpora. Run the three-gate chain (Brin -> skill-auditor -> reputation) before adopting new external packages in production, per AGENTS.md.


Timeline

  • 2026-08-11 | Force-replayed the earlier v2 launch source and complete benchmark image. Added the PyMuPDF-winning counterexample, five-warmup/OCR-off conditions, output-parity boundary, and OCR/LlamaParse escalation replies; current 2.11.1 source remains unchanged. Source: X/@jerryjliu0 v2 launch; LlamaIndex v2 blog; current repository
  • 2026-08-11 | Recovered and inspected all four benchmark screenshots plus the complete browser demo; reframed the headline as a historical disclosed comparison, added explicit competitor and hard-document limits, refreshed all package surfaces to 2.11.1, and registered LiteParse as a stable capability and executable skill object. Source: X/@jerryjliu0 self-thread; GitHub/npm/PyPI; local artifacts, reviewed 2026-08-11
  • 2026-05-27 | @jerryjliu0 (LlamaIndex) announced LiteParse v2 — full Rust rewrite, "world's fastest PDF parser," more accurate than model-free parsers, completely open-source and free. 4,076 likes, 4,561 bookmarks. Source: X/@jerryjliu0, 2026-05-27
  • 2026-06-12 | Dedicated page created from dev-tools bookmark absorption; researched repo (run-llama/liteparse, Apache-2.0, ~9K stars) + LlamaIndex v2.0 blog. Source: x-bookmark-absorb, 2026-06-12
  • 2026-06-30 | One-by-one bookmark artifact review inspected the launch benchmark image, self-thread, OCR reply, upstream repo, npm/PyPI/Git tags, and the upstream Agent Skill. Refreshed routing to make LiteParse the fast local PDF/layout route and imported the skill into skills/engineering/liteparse/. Source: X bookmark artifact audit, 2026-06-30
  • 2026-07-01 | Refreshed package/source snapshot: npm @llamaindex/liteparse and @llamaindex/liteparse-wasm advanced to 2.4.0; GitHub tag docker-v2.4.0 resolves to current HEAD a8288d09cb6bf93c2f7b5257e9634afd26a4566b. Source: npm registry packument; git ls-remote, 2026-07-01
  • 2026-06-30 | Added the later benchmark/demo bookmark: self-thread claims accuracy and speed tests; local video shows the browser parser producing text, Markdown, and JSON from a PDF. Version snapshot was 2.3.0. Source: X/@jerryjliu0, 2026-06-18; Source: npm/PyPI/GitHub, 2026-06-30