Frozen vectors.
Local inference.

Jina v5 text and image embeddings over ONNX Runtime — a Rust core with Python bindings, and one contract every consumer shares.

embroider turns text and images into vectors: task prefix → tokenize → ONNX forward → last-token pooling → L2 → Matryoshka truncate → re-normalise. No torch, transformers, or optimum in the hot path, so every vector in an index stays directly comparable — or the encode fails loudly instead of silently forking the space.

Same bytes in, same vectors out
embed.py Python surface
from embroider import JinaV5

m = JinaV5.open(
  "jinaai/jina-embeddings-v5-text-small-retrieval")
v = m.encode("what do badgers eat?", task="query")
# → 512-dim unit vector, golden-pinned
Load once, encode oftenprefix · pool · truncate · norm
The rule downstream indexes rely on“Never mix what you can't compare.”

Same model, precision, tuning, and length — or re-embed.
Anything else raises instead of drifting.

Illustrative session · weights download once, then reuse the cache.

Keep the space.
Lose the framework.

Embedding pipelines drift when every consumer brings its own tokenizer, pooling, prefix, and runtime. embroider freezes the whole chain in one Rust engine: the task prefix is load-bearing (query ↔ Query:, document ↔ Document:), pooling reads the last attended token, and Matryoshka truncation follows a fixed ladder — then the vector is re-normalised.

Two consumers share it differently. OKFgraph uses the full Jina v5 text + vision contract for its graph. bobine reuses only the ONNX plumbing — session policy, provider fallback, CUDA probe — while owning its own converter-model sessions. Sharing runtime policy never means sharing one global inference session.

  1. PrefixTask decides the space

    Query vs document prefixes, idempotent

  2. EncodeTokenize + ONNX forward

    Session IO discovered at load

  3. PoolLast token, then L2

    int64 ids, float output, clamped index

  4. ShapeTruncate, re-normalise

    Matryoshka ladder, 32–1024 dims

From install
to verified vectors.

One runtime, one lazy session, one assertion. The example below follows the standard path: install the wheel, open the default text model, encode with the task that matches your index, add vision only where a text-nano graph needs it, and verify against the goldens.

Model weights download on first open and reuse the cache afterwards. JinaV5.open is the single expensive step — token counting, search setup, and health checks stay cold until the first real encode.

Install the wheel

Linux / Windows / macOS-arm64 wheels for Python 3.11–3.13. Exactly one ONNX Runtime alongside: CPU or GPU, never both — they share a module name and clobber each other.

pip install embroider
pip install onnxruntime==1.29.0

# Override discovery only when needed:
# export ORT_DYLIB_PATH=/path/to/onnxruntime.dll

Registry at a glance

# small-retrieval (default): fp32 + fp16 mirror
# nano-retrieval: fp32 / fp16 / int8 (vision partner)
# omni-nano vision: fp32 / fp16 → text-nano space
from embroider import available_models
available_models()

Unknown ids fall through to the legacy path: the id is treated as an HF repo holding onnx/model.onnx. Fully local weights skip acquisition via open_files.

1

Open the text model once.

Construction validates the import and device string eagerly; the ONNX session opens lazily on the first encode. device="auto" uses CUDA when the loaded ORT exposes a usable execution provider, else CPU with a warning. precision="auto" follows the landed device — CUDA → FP16, CPU → FP32 — so a missing GPU degrades to FP32 weights instead of stranding FP16 on CPU.

Lazy session
from embroider import JinaV5

m = JinaV5.open(
  "jinaai/jina-embeddings-v5-text-small-retrieval",
  truncate_dim=512, device="auto",
  precision="auto", max_length=None)
2

Encode with the right task.

task is load-bearing: swapped prefixes silently fork the space. Batch encoding stays sequential by design — padded batches waste attention compute on variable-length docs to save milliseconds of boundary overhead. The GIL is released during encode; a failed open is cached and re-raised so config errors fail fast once.

Queries and documents
v  = m.encode("what do badgers eat?", task="query")
vs = m.encode_batch(["a", "b"], task="document")

m.dim()          # effective truncate_dim
m.used_cuda()    # where it actually landed
m.precision()    # effective weights
m.count_tokens("long text...")  # session max_length applies
3

Add vision where it belongs.

Image bytes land in the text-nano space — the registry's text_partner pin refuses anything else. Resize on the caller side to exactly vision_target_size(h, w) with Pillow bicubic; anything else is rejected rather than embedded under the wrong resolution contract. Explicit fp16-on-CPU is an error here: the graph stalls on CPU, so fail-fast beats a hung batch.

Image into text-nano space
from embroider import JinaV5Vision, vision_target_size
from PIL import Image
import numpy as np

v = JinaV5Vision.open()  # omni-nano vision, auto precision
h, w = vision_target_size(img.height, img.width)
rgb = np.asarray(img.resize((w, h), Image.BICUBIC)).tobytes()
vec = v.encode_image(rgb, h, w)
4

Stay cold when you only count.

JinaTokenizer.open fetches only tokenizer.json — about half a second cold, roughly 9× cheaper than a session open — and never truncates, so counts report true length. cache_info answers the pre-flight question — is this model already cached, and where? — offline only: no download, no session, no device probe (precision=None reads fp32). open_files variants skip every download for air-gapped runs; the ONNX sidecar must sit next to the model file, as in the HF cache layout.

Counts and local files
from embroider import JinaTokenizer, cache_info

t = JinaTokenizer.open(
  "jinaai/jina-embeddings-v5-text-small-retrieval")
t.count_tokens("...")  # true length, never truncated

info = cache_info(
  "jinaai/jina-embeddings-v5-text-nano-retrieval",
  precision="fp16")
info["cached"], info["snapshot_path"]

m = JinaV5.open_files("model.onnx", "tokenizer.json")
5

Prove the space, tune the run.

Assert live vectors against the vendored goldens before trusting an index. Then tune the run, not the space: cpu_arena=False cuts peak RSS ~8× for ~1.4× encode time on FP32, and Level3 / intra = physical-cores/2 / inter = 1 stays because it measured fastest. Re-measure on new hardware or ORT before changing policy.

Conformance and tests
cargo test --locked   # 49 pure unit tests, no net/dylib

# vision e2e (ignored): ORT 1.29 dylib + weights + CUDA
ORT_DYLIB_PATH=.../onnxruntime.dll \
  VISION_E2E_FP32=.../model.onnx \
  VISION_E2E_FP16=.../model.onnx \
  VISION_E2E_TOK=.../tokenizer.json \
  VISION_E2E_CUDA=1 \
  cargo test --offline --test vision_e2e -- --ignored

Small surface,
explicit behaviour.

embroider owns session tuning, provider fallback, the CUDA probe, dylib reporting, and owner/name parsing. Consumers own model fetch, preprocessing, IO wiring, and decoding. The seam is deliberate — and documented with a recipe, not just a rule.

Frozen Jina v5 text space

Prefix → tokenize → forward → last-token pooling → L2 → Matryoshka truncate → re-normalise. Session IO is discovered at load; token_type_ids feeds only when declared, output prefers last_hidden_state.

Vision into text-nano space

One image per call over the dynamic-grid graph; host grid tensors are bit-identical to transformers and fixture-pinned. Every encode re-checks seq == image_tokens + 15 so a wrong tokenizer fails loudly.

Precision that follows the device

auto resolves after landing: CUDA → FP16, CPU → FP32. Explicit fp16-on-CPU warns for text and errors for vision. Nano alone offers explicit int8 — never automatic, never mixed into one index.

Memory policy with measurements

CPU arena off by default (15.3 → 1.9 GB peak on FP32 for ~1.4× time). Text tuning is Level3 with intra = physical-cores/2, kept because it measured fastest on a 32-logical-core Windows box.

Reuse without Jina

bobine builds converter sessions from SessionPolicy::ort_defaults() + apply_providers + cuda_available() + report() — never touching JinaV5. Generic fetch helpers remain a natural, additive future.

One shared ORT binary

ort loads dynamically; resolution is ORT_DYLIB_PATH first, else the pip-installed build. OKFgraph resolves it before the native import so ingest and indexing share one binary — no version or CUDA drift.

Fail fast at encode time

No fallback once encoding: vectors must stay bit-comparable within one index. Bad arguments raise before any IO; load failures carry the full anyhow chain. Unknown providers warn and skip; missing CUDA warns and lands on CPU.

Air-gapped and countable

Tokenizer-only opens count exact tokens without a session; open_files paths skip every download; cache_info reports what is already cached, and where — offline only, no session, no device probe. Same bytes in give same vectors out — test-pinned against HF acquisition.

Precision is a
deployment choice.

Weights shift vectors the way tuning levels do — so precision is explicit, reported, and never mixed inside one index. The default follows the machine you actually landed on, not the one you asked for.

On Windows, watch for a stale C:\Windows\System32\onnxruntime.dll (v1.17.x in the wild): with ORT_DYLIB_PATH unset, ort may load it and die with BadVersion. Point the variable at the venv build — the same pitfall bobine documents.

How precision resolves
RequestOn CUDAOn CPUNotes
auto DefaultFP16 weightsFP32 weightsFollows the landed device
fp32FP32 weightsFP32 weightsPinned, always honoured
fp16FP16 weightsFP16 + loud warning (text); error (vision)Default id maps to the FP16 mirror repo
int8Nano only, explicitNano only, explicitNever automatic

Sessions without Jina.

The policy layer is model-agnostic. bobine's engine::session_builder is the reference implementation — tune a builder, apply providers with fallback, commit from file. Inference itself stays with the consumer by design.

// tuned session over your own model file:
let builder = ort::session::Session::builder()?;
let tuned = embroider::SessionPolicy::ort_defaults()
    .apply(builder)?;
let session = embroider::apply_providers(tuned, providers)?
    .commit_from_file(path)?;

// diagnostics for logs:
let rep = embroider::report();
let cuda = embroider::cuda_available();

One engine.
Two languages.

Python covers the Jina contract end to end; Rust exposes the same engine plus the session-policy API other models build on. The pure-Rust crate builds without Python — the extension-module Cargo feature gates PyO3 and is enabled only for wheel builds, the same pattern bobine uses.

Python for embeddings

Open lazily, encode with a task, introspect the landing. Docs live beside the code: quick reference for copy-paste, compatibility matrix for the release story.

from embroider import JinaV5

m = JinaV5.open(
  "jinaai/jina-embeddings-v5-text-nano-retrieval",
  truncate_dim=512)
v = m.encode("hello", task="document")
Quick reference

Rust for policy

Parse device and precision requests, resolve precision against the landed device, parse owner/name ids, and build tuned sessions for any ONNX model — not just Jina.

let device = embroider::DeviceReq::parse("auto")?;
let prec = embroider::Precision::parse("fp16")?;
let landed = prec.resolve(used_cuda);
let policy = embroider::SessionPolicy::text_embed();
API docs

Consumers for graphs

OKFgraph pins the text + vision contract and vendors the goldens; bobine pins the session policy for converter models. The compat matrix records which trio shipped together — and why a newer embroider minor without an OKFgraph update stays supported.

embroider = "0.3"   # shared ORT plumbing
# okfgraph 0.10.x ↔ embroider 0.3.x ↔ ORT 1.29.0
# okfgraph 0.8.x ↔ embroider 0.3.x ↔ ORT 1.29.0
# okfgraph 0.7.x ↔ embroider 0.2.x ↔ ORT 1.29.0
Compatibility

Part of a trio,
usable alone.

embroider ships independently on PyPI and crates.io. It also anchors the embedding side of the OKF stack: OKFgraph owns the knowledge workflow and the embedding-space pins, bobine converts documents — each reusing embroider where it fits, owning its sessions everywhere else.

CPU and GPU runtime distributions are alternatives: install one ONNX Runtime, not both. The wheels cover Linux, Windows, and macOS-arm64 for Python 3.11–3.13.

OKFgraph 0.10.0

Local knowledge graph + agent memory

Website Repository

Uses the full Jina v5 text + vision contract: hybrid retrieval over LadybugDB, model-free PPR when no session is loaded, golden-pinned parity against a numpy replication. The first embedding operation downloads weights; everything cold stays cold until then.

bobine 0.6.0

Documents → Markdown

Website Repository

Reuses the ONNX plumbing — session policy, provider fallback, CUDA probe — for TexTeller, layout, OCR, and table sessions it fetches and owns itself. The recipe is public: any compatible model can follow it with no registry change.

Published stack
ResponsibilityPackageVersion
Knowledge graph · retrieval · MCPokfgraph0.10.0
PDF / Office conversionbobine0.6.0
Jina embeddings · runtime policyembroider0.3.3
Inference runtime (CPU or GPU)onnxruntime1.29.0

Start with one encode
you can trust.

Open the default model, encode a query and a document, assert the goldens. Then hand the vectors to a graph — or the policy to your own model.