Local knowledge.
Connected thinking.

A local knowledge graph and a vault for lasting agent memory. Keep documents, decisions, and working notes searchable, connected, and useful long after the conversation ends.

OKFgraph turns plain Markdown into a local graph with hybrid retrieval. Keep editing in your own tools; use MCP, the CLI, or Python to find the right passage and follow the relationships around it. Embeddings run locally over ONNX Runtime, without torch.

Knowledge that outlives the session
kb/auth.md Markdown source
---
title: Session auth
tags: [auth, sessions]
---

## Design decision
Use short-lived access tokens.
Refresh flow: [[token-refresh]]
Index locallychunks · vectors · links
Recall in the next session“Why did we choose this auth design?”

Find the decision. Read the reasoning.
Resume with the context intact.

Illustrative workflow · your files remain yours.

Keep the files.
Gain the connections.

A folder of notes is easy to own, but harder to query as it grows. Vector search finds related text, but doesn't necessarily explain how one decision connects to another. OKFgraph brings these views together: your Markdown bundle supplies the content, and a LadybugDB index stores concepts, passages, links, vectors, and full-text indexes.

In a file-backed workflow, edit and review the bundle as normal, then import the changes. You can rebuild the index from those sources or export graph content back to Markdown. There is no hosted service to run, and no proprietary authoring format to adopt.

  1. SourceMarkdown bundle

    Notes, metadata, links, and assets

  2. IndexImport changes

    Parse, chunk, embed, resolve links

  3. RetrieveLadybugDB graph

    Combine meaning, words, and structure

  4. UseRead or export

    Relevant context for people and agents

From a first note
to a working memory.

Start with one folder and one local index. The example below follows a small authentication knowledge base through its full lifecycle: write the source, import it, retrieve context, record a new decision, and check the graph before sharing it.

Run the commands from your project directory. Configuration keeps subsequent calls pointed at the same database and bundle; the first embedding operation downloads model weights, and later calls reuse the cache.

Install and configure

Use the CPU extra for local embeddings. Add pdf only when you need PDF or Office conversion; use the GPU extra instead of CPU for a CUDA setup, not alongside it.

pip install "okfgraph[cpu]"
# Optional document conversion:
# pip install "okfgraph[cpu,pdf]"

Save as okfgraph.toml

bundle_root = "kb"

[database]
db_path = "kb.db"

[embedding]
embedding_dim = 512
device = "cpu"

init initializes the database; it does not write this configuration for you.

1

Write a note you can own.

Create kb/ and save auth.md inside it. Frontmatter describes the note; its relative path gives it the concept ID auth. Save a second note as token-refresh.md to make the wikilink resolvable. Your editor and git remain the place to author and review knowledge.

kb/auth.md
---
title: Session auth
tags: [auth, sessions]
---

## Design
Use short-lived access tokens.
See [[token-refresh]] for renewal.
2

Build the index, then update it.

Import parses the notes, creates searchable chunks, embeds the content, and resolves links. Content hashes let later imports skip unchanged files. Edit a decision in your bundle and run import again; you don't need to re-embed the entire corpus each time.

Initialize and import
okf init
okf import --all

# After editing your notes:
okf import --all

# Explicitly remove deleted sources:
# okf import --all --prune-missing
3

Find the passage, follow the context.

Ask an open question with hybrid search, or use PPR for lexical-seeded graph retrieval without loading an embedding model. Read the selected concept with a token budget and surrounding context, then traverse links when you need to inspect the relationships rather than just the prose.

Search, read, traverse
okf search "how are sessions renewed?"
okf search "auth" --rank ppr
okf read auth --include context \
  --max-tokens 1500
okf traverse auth --direction BOTH
4

Keep the reasoning, not just the result.

Record a decision before the session ends. Thoughts become searchable concepts under IDs such as thoughts/auth/…, without needing a source file. For a paper or Office document, install the PDF extra and let bobine stage the conversion to Markdown before importing it.

Add working knowledge
okf ingest --kind thoughts \
  --thoughts "Use rotating refresh tokens." \
  --topic auth

okf ingest --kind pdf \
  --pdf-path paper.pdf
5

Check the graph. Share the knowledge.

Doctor reports graph health; diff shows drift between the index and its source folder. Use strict checks in CI to make findings visible. Export writes graph content to a new directory, including thoughts created without files, with an Obsidian flavor when you want a vault.

Maintain and export
okf doctor --strict
okf diff
okf export --all --output-dir out/ \
  --flavor obsidian

Useful at every stage
of the knowledge loop.

OKFgraph combines retrieval with the maintenance work a long-lived corpus needs. These capabilities share the same local graph, rather than introducing a separate service for every step.

Search meaning and exact words

Jina v5 vectors and full-text matches are fused with reciprocal rank fusion (RRF). Search concepts for an overview or chunks for a specific passage. Matryoshka dimensions let you choose a smaller vector representation when creating your index.

Use the graph without a model

PPR starts from lexical matches and expands through links with Personalized PageRank. It is useful for topic exploration and cold sessions: no query embedding or ONNX session is needed for this retrieval mode.

Import only what changed

Delta-aware import avoids re-embedding unchanged files. Multiple source roots can share one graph under @alias/ namespaces; deletion tracking distinguishes a missing source from an unmounted root.

Make corpus health visible

Doctor reports broken links, orphans, stale content, duplicates, and missing descriptions with a 0–100 score; model and converter cache status appears as unscored info. Diff exposes source/index drift. Safe repairs and structured output support repeatable maintenance and CI checks.

Work with an Obsidian vault

Wikilinks resolve through stable IDs, aliases, titles, and file stems; ambiguous names remain unresolved rather than being guessed. The Obsidian export flavor preserves graph links for vault-oriented workflows.

Retrieve image content

Use captions in text-only mode, or ONNX vision embeddings for image-aware retrieval. Vision uses omni-nano paired with a text-nano graph, with content-hash assets and an explicit preprocessing contract.

Give agents a lasting memory vault

Thought ingestion stores decisions and reasoning as concepts with topic-based namespaces. Search, read, and traverse recover them across sessions, beyond a single context window; export materializes them as Markdown when you want a file-backed record.

Turn structured data into notes

okf produce --from sqlite creates one concept per table, links foreign keys, and adds observations about missing values, duplicates, and orphan references. A lint pre-flight checks the bundle before you import it; log.md records producer changes.

Choose the signal
your question needs.

Use hybrid retrieval for open-ended questions, hub ranking when well-linked concepts should carry more weight, and PPR when relationships are the useful signal. The default preserves RRF ranking; the other modes are explicit choices, not hidden re-ranking steps.

Fixed graph-ranking rules, stable ordering, and checked-in fixtures support reproducible tests. Embedding conformance uses numerical tolerances: changing hardware, precision, or model weights is not a promise of bit-identical results. Model and precision pins help prevent incompatible embedding spaces from being mixed in one index.

Three retrieval modes over one graph
ModeBest fitQuery embeddingSignal
rank=none DefaultOpen-ended questionsYesVector + full-text, fused with RRF
rank=hubLink-authority weightingYesHybrid ranking blended with hub scores
rank=pprTopic exploration, cold sessionsNoLexical seeds + Personalized PageRank

Small pieces. Explicit contracts.

The implementation separates conversion, embedding, import, retrieval, and export behind the router facade. Rust handles document conversion and ONNX inference; LadybugDB holds the graph, vectors, and full-text indexes. That division keeps responsibilities inspectable without requiring a fleet of services.

# A graph query without an embedding session:
okf search "auth" --rank ppr

# Retrieve passages rather than whole concepts:
okf search "token renewal" --target chunks

# Read with a bounded context budget:
okf read auth --include context \
  --max-tokens 1500

One graph.
Three ways in.

MCP gives agents a small tool surface, the CLI supports interactive work and automation, and Python exposes the router for your own pipelines. All use the same underlying components. Bundle import and maintenance stay outside the eight MCP tools, so agents don't need a large administrative tool catalogue; anything snapshot-shaped (diff, doctor, lint, produce) is CLI-only.

MCP for agents

Connect with stdio. The eight tools are search, read, traverse, ingest, export_bundle, export_concept, list_images, and get_image. Every tool returns the result envelope; failures carry isError with the typed error. Use persistent, absolute paths appropriate to your machine.

{
  "mcpServers": {
    "okfgraph": {
      "command": "okf-mcp",
      "args": [
        "--db-path", "/path/to/kb.db",
        "--bundle-root", "/path/to/kb"
      ]
    }
  }
}

CLI for people and CI

Use the same retrieval verbs alongside import, diff, doctor, and export. Each command has help; configuration, environment variables, and flags let you choose how explicit to be.

okf search "token renewal"
okf read auth --include context
okf traverse auth --direction BOTH

# Inspect drift and graph health:
okf diff
okf doctor --strict
CLI reference

Python for pipelines

Import a bundle, retrieve concepts, and export directly through OKFRouter. Omit bundle_root for graph-only thought workflows; close the router when you're done.

from pathlib import Path
from okfgraph import OKFRouter

r = OKFRouter(bundle_root="./kb",
              db_path="kb.db")
try:
    r.import_bundle(None)
    hits = r.search("auth", limit=5)
    r.export_bundle(Path("out"))
finally:
    r.close()

A focused stack,
with clear ownership.

OKFgraph owns the knowledge workflow and the embedding-space pins. Two independently usable Rust-backed libraries support it: bobine converts documents, and embroider supplies Jina embeddings and shared ONNX session policy. Each consumer owns its sessions; sharing runtime policy doesn't mean sharing one global inference session.

The graph store is LadybugDB, not SQLite. SQLite is supported as a source for bundle production. CPU and GPU runtime extras are alternatives: install one ONNX Runtime distribution, not both.

bobine 0.6.0

Documents → Markdown

Website Repository

PDF and Office conversion stays behind a replaceable converter boundary. PDF routes let you choose text extraction, surgical formula processing, or full layout/OCR analysis. DOCX, XLSX, PPTX and their legacy counterparts enter through extension-based dispatch; Excel also offers per-sheet CSV and typed JSON output.

Provider defaults follow measurement: CUDA helps layout and OCR, while SLANet-plus tables remain CPU-pinned. TexTeller defaults to int8 CPU, with fp32 CUDA available as an opt-in. Faster hardware is only useful when the model's graph benefits from it.

Inspect the CPU / CUDA measurements

RTX 3090 results from the published bobine benchmarks, measured with ORT 1.28.1 before the current 1.29.0 pin. These are workload-specific observations, not latency guarantees for your machine.

Published conversion benchmarks
WorkloadCPUCUDARouting choice
Layout 1024²480 ms37 msCUDA when available
OCR det + rec 1024²131 ms36 msCUDA when available
OCR 47-line page0.79 s0.32 s (batched ×32)CUDA + batching
Table 700 × 40028.8 ms49 msCPU pinned
Table 1024²133 ms893 msCPU pinned
TexTeller int80.50 s1.32 sCPU default
TexTeller fp320.95 s0.42 sCUDA opt-in
Methodology and full results

embroider 0.3.3

Jina v5 → ONNX embeddings

Website Repository

embroider combines a Rust embedding engine with Python bindings and an ONNX session-policy API. OKFgraph uses its Jina v5 text and vision encoders; bobine uses the runtime policy for its own conversion models, not Jina inference.

Text pooling, normalization, model dimensions, and vision preprocessing are explicit contracts. Checked-in conformance fixtures test numerical alignment. Automatic precision follows the device: fp32 CPU or fp16 CUDA; vision fp16 on CPU is rejected. Text models also offer an explicit int8 option.

API quick reference · Runtime compatibility

Published stack
ResponsibilityPackageVersion
Knowledge graph · retrieval · MCPokfgraph0.10.0
PDF / Office conversionbobine0.6.0
Jina embeddings · runtime policyembroider0.3.3
Graph · vectors · full-text indexladybug0.21.2
Inference runtime (CPU or GPU)onnxruntime1.29.0

Start with the knowledge
you already have.

One local graph for the documents you own and the memory your agents keep. Build a small corpus first, save a useful decision, and pick it up again in the next session.