Open the text model once.
Construction validates the import and device string eagerly; the ONNX session opens lazily on the first encode. device="auto" uses CUDA when the loaded ORT exposes a usable execution provider, else CPU with a warning. precision="auto" follows the landed device — CUDA → FP16, CPU → FP32 — so a missing GPU degrades to FP32 weights instead of stranding FP16 on CPU.
from embroider import JinaV5
m = JinaV5.open(
"jinaai/jina-embeddings-v5-text-small-retrieval",
truncate_dim=512, device="auto",
precision="auto", max_length=None)