Documents in.
Markdown out.

A standalone PDF, Office, and text ingestion engine — a pure-Rust core with Python bindings, and no ML framework in the import path.

bobine runs on a single onnxruntime shared library with no CUDA-version coupling: pdf_oxide for fast native text extraction, ONNX models for the heavy passes — TexTeller formula OCR, DocLayout-YOLO layout, PaddleOCR. No torch, no optimum, no opencv — not even on the Python side.

One converter, four routing modes
convert.py Python surface
import bobine

conv = bobine.HybridConverter(
  bobine.ConverterConfig(
    routing_mode=bobine.RoutingMode.Surgical),
  cache_dir="~/.cache/huggingface/hub")
md = conv.convert_pdf("paper.pdf", work_dir="/tmp/out")
# → Markdown with $…$ / $$…$$ LaTeX in place
Fast path firstheavy passes only where flagged
The rule the router keeps“Never fail a conversion over a model.”

Missing or broken heavy models degrade to the fast path per page.
Text and Office work needs no models at all.

Illustrative session · converter weights auto-download into the shared hub cache.

Keep the pipeline.
Drop the framework.

The ingestion pipeline was originally entangled with the knowledge-graph project it served (OKFgraph). bobine moves conversion into its own package, so any consumer — a graph, a CLI, an MCP server, a batch tool — can reuse it without importing a database stack.

The v0.3.0 rewrite ports the whole pipeline to Rust: same routing heuristics and output contract as the proven Python implementation (preserved under legacy/), with native speed and no Python ML dependencies. Dual-licensed MIT OR Apache-2.0 — pick either.

  1. SourcePDF / Office / text

    Extension dispatch, binary-sniffed text

  2. RouteHeuristics per page

    Math signal, scanned detection

  3. ConvertFast path + ONNX passes

    Formulas, layout, OCR, tables

  4. EmitMarkdown + assets

    LaTeX in place, staged figures

From a PDF
to trusted Markdown.

One converter covers papers, scans, Office files, and plain text. The example below follows the standard path: install the wheel with one runtime, convert surgically, force the full pipeline only for scans, recognise a lone formula crop, and stage a document-level ingest.

Converter weights auto-download from HuggingFace on first use into the shared hub cache (`~/.cache/huggingface/hub` unless `cache_dir` points elsewhere). Office, text, and fast-path PDF work needs no models at all.

Install with one runtime

Wheels for Linux, Windows x86_64, and macOS arm64. The cpu and gpu extras are alternatives — never install both onnxruntime packages. import bobine locates the library; an already-set ORT_DYLIB_PATH always wins.

pip install bobine
pip install "bobine[cpu]"
# or: pip install "bobine[gpu]"  # self-contained CUDA

# Rust library (no Python involved):
# cargo add bobine

Converter weights

# TexTeller int8 default  ~319 MB  (Ji-Ha export)
# TexTeller fp32 opt-in    ~1.25 GB (OleehyO)
# DocLayout-YOLO           ~72 MB
# PP-OCRv4 det + rec       ~16 MB
# SLANet-plus tables       ~7.8 MB

Missing layout/OCR models degrade to the fast path per page — a conversion never fails because of them. A missing table model only disables scanned-table recognition.

1

Convert a paper surgically.

Surgical keeps born-digital pages on the fast path and spends TexTeller only on formula crops — the economical default for papers with math. Formula regions come from the PDF text layer (TeX math fonts plus unicode math), merged line-aware so multi-line display equations become one crop while columns and prose stay apart.

Surgical conversion
import bobine

conv = bobine.HybridConverter(
  bobine.ConverterConfig(
    routing_mode=bobine.RoutingMode.Surgical),
  cache_dir="~/.cache/huggingface/hub")
md = conv.convert_pdf("paper.pdf", work_dir="/tmp/out")
2

Force the full pass for scans.

Always runs every page through layout + OCR. On born-digital pages, formula regions are refined against text-layer math boxes (budgeted TexTeller decode + plausibility filter), so output matches Surgical while scans stay on the full crop path. For text-layer-hostile PDFs, add formula_layout_fallback=True to ask the layout model for equation regions.

Full pipeline and fallback
# scanned document — full pipeline at high dpi
cfg = bobine.ConverterConfig(
  routing_mode=bobine.RoutingMode.Always, render_dpi=300)
conv = bobine.HybridConverter(cfg, cache_dir="~/.cache/huggingface/hub")
md = conv.convert_pdf("scan.pdf", work_dir="/tmp/out")

# fastest: text layer only, never touch ONNX
fast = bobine.ConverterConfig(
  routing_mode=bobine.RoutingMode.Never, use_onnx=False)
3

Recognise one formula.

TexTeller decodes a ViT encoder → RoBERTa decoder ONNX graph with KV-cache — roughly 5× faster per crop than the legacy RapidLaTeXOCR backend, with markedly better accuracy. Preprocessing matches upstream exactly, including its normalize-before-pad order. On a 10-formula corpus, Int8+CPU scores 10/10 and Fp32+CUDA 9/10 — single-token greedy-decode noise, not systematic damage.

Formula crop to LaTeX
latex = conv.recognize_formula("crop.png")
# r"\frac{1}{2}"

# one-shot helper for whole files:
md = bobine.convert_to_markdown(
  "paper.pdf", bobine.ConverterConfig(),
  work_dir="/tmp/out", cache_dir="~/.cache/huggingface/hub")

# text needs no models at all:
md = conv.convert("notes.txt", work_dir="/tmp/out")
4

Bring Office along.

.docx, .xlsx, .pptx and legacy .doc / .xls / .ppt convert via office_oxide — no models needed. Embedded pictures stage into <work_dir>/assets/office/ under content-hash names; workbooks additionally export per-sheet CSV and typed JSON siblings.

Office and workbooks
md = conv.convert("deck.pptx", work_dir="/tmp/out")

doc = bobine.convert_excel("book.xlsx")
doc.markdown       # ## {sheet} sections + GFM tables
doc.sheet_names    # ['Data', 'Summary']
doc.csv_by_sheet   # {'Data': 'Item,Count,...\n', ...}
5

Stage document-level ingest.

ingest_document and convert_directory return versioned ConvertedDocuments with frontmatter and lint hooks — the shape graph pipelines consume. OCR recognition batches line crops in chunks of 32 on accelerators (6.3× faster on CUDA) while CPU sessions keep the exact single-line path with byte-identical tensors.

Ingest with hooks
bobine.ingest_document(
  "paper.pdf", "out/",
  lint_callback=lambda md: (True, md.strip() + "\n"),
  on_page=lambda i, n: print(f"page {i+1}/{n}"),
  should_continue=lambda: not cancelled())

Useful at every stage
of the ingest loop.

bobine owns fetching, preprocessing, session wiring, and decoding for every converter model. embroider contributes only the session policy, provider fallback, and CUDA probe — it never sees converter weights. That division keeps responsibilities inspectable. model_status() reports which converter weights are already cached — offline, per family.

Four routing modes

Never for the pure fast path, Auto for per-page heuristics, Surgical for TexTeller-only formulas, Always for full layout + OCR with hybrid formula refinement on born-digital pages.

TexTeller formula OCR

Int8 by default (~319 MB, 10/10 on the corpus, ~0.5 s median on CPU); fp32 opt-in (~1.25 GB, ~0.42 s on CUDA). Int8 + CUDA is discouraged and warned against — quantized ops lose on the GPU path.

Layout, OCR, tables

DocLayout-YOLO page analysis, PaddleOCR det + rec with provider-gated batching, SLANet-plus tables converted from HTML to GFM pipe tables. Tables stay CPU-pinned by measurement, not habit.

Office without models

DOCX, XLSX, PPTX plus legacy DOC / XLS / PPT through extension-based dispatch. Excel yields per-sheet CSV siblings and typed JSON; pictures stage under content-hash names.

Output you can diff

One Markdown string: inline $…$ and display $$…$$ LaTeX spliced in place, GFM tables, fenced code from monospaced runs, embedded images in work_dir, plus an okf-asset:// staging store.

Graceful under failure

Heavy models degrade to the fast path per page; binary files under text-ish extensions fail fast with UnsupportedFormat instead of decode garbage. A conversion never fails because a model did.

GPU with zero config

Layout and OCR auto-enable CUDA when the loaded library exposes a usable EP (shared probe since v0.5.10). Per-slot *_ort_providers overrides pin any slot; CoreML entries cover Apple Silicon.

Tested against fixtures

137 lib tests plus office, excel, and golden suites; trimmed CC BY 4.0 arXiv papers with provenance in SOURCES.md. The frozen Python v0.2.0 tree stays under legacy/ as the reference implementation.

Measured speed,
per workload.

Reference box: RTX 3090, ORT 1.28.1 — the benches predate the 1.29.0 pin, so read ratios, not absolutes. Full tables, fixtures, and repro commands live in the benchmarks doc.

Faster hardware only helps when the model's graph benefits from it. Quantization is hardware-dependent; batching is provider-gated. The router encodes those measurements as defaults.

Published conversion benchmarks
WorkloadCPUCUDARouting choice
Layout 1024²480 ms37 msCUDA when available
OCR det + rec 1024²131 ms36 msCUDA when available
OCR 47-line page0.79 s0.32 s (batched ×32)CUDA + batching
Table 700 × 40028.8 ms49 msCPU pinned
Table 1024²133 ms893 msCPU pinned
TexTeller int80.50 s1.32 sCPU default
TexTeller fp320.95 s0.42 sCUDA opt-in

Rust for pipelines.

The same converter serves both languages. Rust owns the pipeline directly; Python drives it through PyO3 behind the extension-module feature, while the crates.io build stays Python-free.

use bobine::{ConverterConfig, HybridConverter, RoutingMode};

let mut conv = HybridConverter::new(
    ConverterConfig {
        routing_mode: RoutingMode::Surgical,
        ..Default::default()
    },
    std::path::Path::new("~/.cache/huggingface/hub"),
);
let md = conv.convert_pdf(
    std::path::Path::new("paper.pdf"),
    std::path::Path::new("/tmp/out"))?;

One engine.
Three doors.

Python covers the converter, one-shot helpers, and staged ingest; Rust exposes the same pipeline natively; document-level workflows return versioned documents with frontmatter for graph pipelines. Text and Office paths never load a model.

Python for conversion

Direct converter, one-shot Markdown, formula crops, and staged ingest with lint, progress, and cancellation hooks.

conv.convert_pdf("paper.pdf",
                 work_dir="/tmp/out")
conv.recognize_formula("crop.png")
bobine.convert_to_markdown(
  "paper.pdf", config,
  work_dir="/tmp/out",
  cache_dir="~/.cache/huggingface/hub")
Quick reference

Rust for embedding

Add the crate, build the converter, route natively. Sessions are built on embroider's policy with bobine-owned fetch and decoding.

cargo add bobine

BobineConverter::default()
// + ConverterConfig knobs:
// routing, dpi, providers,
// precision, quantization
API docs

Graphs for context

OKFgraph consumes staged Markdown behind a replaceable converter boundary; embroider supplies the session policy bobine builds on. Each package versions and ships alone.

# one shared ORT binary:
onnxruntime==1.29.0
# bobine 0.6.x ↔ embroider >=0.3.3 (policy + cache_info_files)
# okfgraph 0.10.x ↔ bobine 0.6.x
Architecture

A focused trio,
with clear ownership.

bobine owns document conversion end to end. OKFgraph owns the knowledge workflow that consumes its Markdown; embroider owns the session policy its ONNX sessions are built on. Each consumer owns its sessions; sharing policy doesn't mean sharing inference state.

The v0.6.0 stack shares one ONNX Runtime 1.29.0 binary across all three packages, located via ORT_DYLIB_PATH or the pip-installed build. On Windows, a stale System32\onnxruntime.dll (1.17.x) breaks the load — point at the venv build.

OKFgraph 0.10.0

Local knowledge graph + agent memory

Website Repository

Consumes staged Markdown behind the DocumentConverter seam — install the PDF extra and bobine stages the conversion before import. Bundle import and maintenance stay outside the agent tool surface.

Inspect the CPU / CUDA measurements

RTX 3090 results from the published bobine benchmarks, measured with ORT 1.28.1 before the current 1.29.0 pin. These are workload-specific observations, not latency guarantees for your machine.

Workload-specific routing evidence
WorkloadCPUCUDATakeaway
Layout 1024²480 ms37 ms (~12.8×)CUDA auto-enabled
OCR det + rec131 ms36 ms (~3.6×)CUDA auto-enabled
OCR batched ×324.8× slower6.3× fasterProvider-gated batching
Tables28.8–133 ms2–9× slowerCPU pinned
TexTeller int80.50 s1.32 sCPU default
TexTeller fp320.95 s0.42 sCUDA opt-in
Methodology and full results

embroider 0.3.3

Jina v5 → ONNX embeddings

Website Repository

Supplies the session policy (SessionPolicy::ort_defaults()), provider fallback, and CUDA probe bobine builds converter sessions from. Converter weights, preprocessing, and decoding stay with bobine by design. Its shared cache_info_files lookup also powers model_status(): one offline cache report per converter family.

API quick reference · Runtime compatibility

Published stack
ResponsibilityPackageVersion
Knowledge graph · retrieval · MCPokfgraph0.10.0
PDF / Office conversionbobine0.6.0
Jina embeddings · runtime policyembroider0.3.3
Inference runtime (CPU or GPU)onnxruntime1.29.0

Start with the document
you already have.

One converter for the papers, decks, sheets, and notes you own. Convert surgically first, go full-pipeline for scans, and hand the Markdown to a graph.