← Blog · Download & Install · Models

Document parsing on CPU: page images to structured Markdown with Nemotron-Parse 2.0

September 2026 · SearchAI Inference Server team

new in v1.4.0sub-1B document OCR Q8 · ~1 GB filefits 8 GB hosts CPU-only

A lot of document work never needed a general-purpose vision model. It needed one job done reliably: take the picture of a page — a scan, a screenshot, a rendered PDF page — and turn it into structured text you can actually process. Headings, paragraphs, tables, and where each of them sits on the page.

As of v1.4.0, the SearchAI Inference Server natively serves Nemotron-Parse 2.0 — NVIDIA's sub-1B document-OCR model. A page image goes in; structured Markdown comes out: text with layout classes and bounding boxes. It runs CPU-only, on an 8 GB host, from a single model file — and it speaks the same OpenAI-compatible API as everything else on the server, so it's the "model" field on a chat request.

What's inside the model

Nemotron-Parse 2.0 is a compact document-understanding pipeline:

  • a C-RADIO ViT-H/16 vision trunk that reads the page image,
  • a conv neck that projects the trunk's features, and
  • a 10-layer mBART decoder that emits the structured text.

We serve it as Q8 — a ~1 GB model file — which is why it fits comfortably on an 8 GB machine. There is no separate mmproj file: the vision trunk lives inside the single gguf, so fetching one file gives you the whole image-to-Markdown pipeline.

The API is just chat completions

There's no new endpoint to learn. You call /v1/chat/completions with "model": "nemotron-parse-2.0", pass the page as an image content part — a base64 data URL or a public https URL — and add a short instruction. The assistant message's content is the structured Markdown.

curl -s http://127.0.0.1:8081/v1/chat/completions \
  -H "Authorization: Bearer $SEARCHAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "nemotron-parse-2.0",
    "messages": [{
      "role": "user",
      "content": [
        {"type": "text", "text": "Parse this document."},
        {"type": "image_url",
         "image_url": {"url": "data:image/png;base64,iVBORw0KGgoAAAANSUhEUg..."}}
      ]
    }]
  }'

The image can equally be a public URL — "url": "https://example.com/page-01.png" — if the server can reach it. Either way, the instruction stays short: the model's job is the parse, not a conversation.

What it returns

Here's a validated example. We handed it a chart image — a simple bar chart with four categories — and it came back as a clean Markdown table, no post-processing:

| Category | Value |
| --- | --- |
| Red | 100 |
| Blue | 85 |
| Green | 70 |
| Purple | 60 |

That's the shape of the output across a page: text carrying its layout — headings, paragraphs, list items, table cells — each with a class and a bounding box, assembled into Markdown you can index, diff, or feed to another model. The same request produced byte-consistent output on both arm64 and amd64, so a parse run on a Graviton box and the same parse on an Intel or AMD box agree exactly.

Performance: it scales with cores

The honest performance story is that the C-RADIO ViT trunk is compute-bound. Reading the page image is the expensive part, and that work parallelizes across CPU cores. So latency scales with core count: on a many-core server the parse is fast; on a small box it's slower. If throughput matters, the lever is cores, not clock.

There's no separate mmproj to manage and no GPU requirement — the vision trunk is inside the single gguf and the whole thing runs on CPU. A many-core server is where it's happiest.

Getting it

Install the server, then fetch the nemotron-parse model group (or load it from the console's Models tab). It then serves alongside your existing models from the same process.

curl -fsSL https://inference-server.searchblox.com/install | sudo bash
sudo ./fetch-models.sh nemotron-parse
# then call it with: "model": "nemotron-parse-2.0"

Nemotron-Parse 2.0 is NVIDIA's sub-1B document-OCR model, served here as Q8 (~1 GB) and running CPU-only on an 8 GB host. The chart-to-table example above is a real, unedited parse; output was verified byte-consistent across arm64 and amd64.