Teach a model
what you know.

Your knowledge as a searchable corpus, a stateless model inside an agent loop, and two kinds of training that make both sides better. Called like any language model.

01 · Corpus, chunks, embeddings

From documents
to vectors.

The knowledge stays text. To find the right piece later, every piece is cut to a size one embedding can represent and turned into a vector.

  1. importPDF, Word, Markdown, text · CSV, Excel
  2. sections~2,400 chars, split at paragraphs, 200 chars overlap
  3. unitsAssist extracts typed units; table cells become facts without a model
  4. chunksunits over 4,000 chars split to ~1,600: paragraph → sentence → hard cut
  5. embedQwen3-Embedding-0.6B, 1024 dims
  6. indexone float32 matrix in memory + a BM25 index
a unit and what gets embedded
{ "type": "fact", "title": "Module X buffer size", "content": "Module X uses a 256-byte receive buffer.", "meta_json": { "subject": "Module X", "property": "Buffer size", "value": "256", "unit": "byte" }, "validated": false, # proposed, works, not yet accepted "core_marking": "tail"}embedding_text "Module X · Buffer size: 256 byte"query prefix "Instruct: Given a user question, retrieve knowledge passages that answer it\nQuery: "
why chunkone vector for a long text is an average of everything in it; a manual of 87 KB embedded as one unit matched nothing well
what is embeddedtitle + content; a fact as its fact line; a case as its situation and criteria, the problem rather than the write-up
search speedone matrix-vector product: about 10 ms at 50,000 units
calibrationraw cosine is mapped per agent so the best relevant/irrelevant separator lands at 0.30 and the median relevant hit at 0.60, whatever the embedding model

Six unit types: fact concept principle skill process case, linked by 14 edge types such as causes, requires, replaces or contradiction.

02 · Retrieval

The model can't read everything.
So it gets what fits the question.

A context window holds a few thousand words of knowledge, a corpus holds millions. Retrieval picks the handful of units that answer this question, so the model works from your text instead of from what it half remembers.

  1. understandquestion type: lookup · explain · diagnose · advise · instruct · review
  2. densemeaning: cosine over the vectorsBM25exact words, codes, numbers
  3. RRFΣ 1/(60 + rank) merges both rankings
  4. rerankoptional cross-encoder over 30 candidates
  5. top 5score ≥ 0.25, until the knowledge budget is full
  6. graph≤ 2 hops along the edges the question type needs
gate
top_score < 0.30 → refusal

If nothing scores high enough and there is no core knowledge or tool to answer from, the model is not even asked.

core knowledge
core: always in the prompt

The few rules and concepts every answer needs skip retrieval and sit in the system prompt, inside their own token budget.

graph walk
diagnose → causes, mitigates, requires

A diagnosis pulls in causes and fixes, an instruction its prerequisites. An outdated unit whose replacement was found drops to the end.

03 · The model in the agent block

The model is stateless.
The request is not.

A language model remembers nothing between calls. Nemonix gives one request a state of its own: the retrieved knowledge, the tool results gathered so far and a round counter. When the answer is out, that state is gone and the next request starts clean.

one request · example
> is the H7 board still within its licence?state knowledge: 4 units · tools: [] · round 0/4round 1 model → licence_status(board="H7") mcp server (hidden) → expires 2027-03-01round 2 model → answercheck claims=2 removed=0sources Licence terms for boardsdone state discarded
tool loopup to 4 rounds (1–20); when they are used up, one last call without tools forces the answer. The loop is bounded by Nemonix, not by the model
hidden toolsMCP servers you attach to the expert, such as analysis or lookup tools. Callers only ask questions; they never see, provide or reach these tools
when tools runper tool: freely in the loop, only for certain question types, or always once before answering
tool resultsenter the context as fenced data with their own token budget, never as instructions
caller historythe API is stateless: the caller sends the conversation with each request
The context of one request. Pick a model's window:
knowledge
12 % of the window, at least 4,000 tokens

The units retrieved for this question, under “## VERIFIED KNOWLEDGE”, “## EXPERIENCE CASES” and “## RELATED KNOWLEDGE”. Hits are added in score order until this budget is full.

After the model answers, a second call checks it claim by claim against the knowledge and removes what isn't covered; sources are then attached without a model.

04 · Training the retriever

Teach the search
your vocabulary.

A general embedding model doesn't know that two of your terms mean the same thing. From 100 units on, it can be fine-tuned on your own corpus in the Finetune Hub, on your own GPU.

data
[question, unit, hard negative]

Questions come from a generator, from people and from your test set. The hard negative is the nearest unit that has no edge to the right one.

loss
MultipleNegativesRankingLoss

Contrastive: the right unit has to beat its hard negative and every other unit in the batch. One epoch, batch 16, 20 % held out.

adoption
recall@5 ↑ and nDCG@5 ≥

Otherwise the old model stays. Adopting re-embeds the corpus and recalibrates the scores.

0.827 → 0.899*
recall@1
0.887 → 0.934*
MRR
6 GB
GPU (RTX 2060)
0.6 B
parameters

* Internal measurement on one test corpus, July 2026. Your numbers depend on your knowledge.

05 · Adapter fine-tuning

Train how it uses knowledge,
not the knowledge itself.

A full fine-tune rewrites every weight, forgets more and leaves one large model per customer. A LoRA adapter trains about 2 % of the parameters on a frozen base; with QLoRA that base is held in 4 bit. It learns less, and that is fine: the facts it should apply arrive in the context with every question, through retrieval and core knowledge. Trained in the Finetune Hub, on your own GPU.

one training example
systemAnswer from the given knowledge only. If it does notcover the question, say so plainly instead of guessing.Copy values exactly. Do not cite sources.userKNOWLEDGE:[1] Module X uses a 512-byte receive buffer.[2] … [3] … [4] … # the right unit + noiseQUESTION: receive buffer size of module X?assistantModule X uses a 512-byte receive buffer.
methodLoRA, or QLoRA with the base in 4-bit NF4; adapters on every linear layer, rank 16
base modelyour choice: any open model you can run
what it learnsto read the given knowledge, copy values exactly, reason the way your expert does, and refuse what the context doesn't cover
counterfactualsexamples change a known value in the context, so the model learns to trust the text over its memory
costan adapter is tens to low hundreds of MB over one shared base, instead of a full model copy per customer (~54 GB at 27B)
memoryestimated ~25 GB to train a 27B base, ~10 GB for 8B
06 · Assist

An AI that helps you
build the expert.

Assist is an agent of its own inside the Studio. It reads what is already there, proposes new knowledge, and tests the expert on a sample question, while your expert stays the one who accepts.

assist · tools
db_lookup_similar is this already known?db_lookup_topics where does it belong?check_conflict does it contradict a unit?extract_core_materialtest_pipeline run the expert on a questionup to 8 tool calls per turn/case-interview context · decision · rationale · outcome · criteria/import documents into units
proposestyped units with synonyms and a core suggestion, topics, edits and archivals, and pipeline changes
case interviewasks for a real decision in five parts, the way experts recall critical decisions
where it landsthe Inbox: written at once, marked validated: false, accepted by a person
modelyour own key, or managed: Claude Haiku from Solo, Sonnet from Team, Opus from Business

Outside agents can propose knowledge too, over MCP. They land in the same Inbox and can't validate anything themselves.

Modules

The Studio builds.
The modules run it.

Nemonix Studio

The desktop app: knowledge, pipeline, Playground, test questions, releases.

Nemonix Serve

Always-on runtime for released versions. A version freezes knowledge, embeddings, model and pipeline.

Nemonix GPU Engine

Local embeddings on your own GPU, from 6 GB of VRAM.

Nemonix Central

A private workspace server for your team, on any host in your network.

Nemonix Finetune Hub

Fine-tune the retriever and the LLM, on your own hardware.

API

Called like any model.

Serve speaks the OpenAI and Anthropic formats. Claude Code or OpenCode can use an expert as their model; their own tools are chosen by Serve but run on the caller's side, never inside the expert.

POST /v1/agents/query
{ "prompt": "receive buffer size of module X?" }→ 200{ "answer": "Module X uses a 256-byte receive buffer. \n\nSources:\n- Module X buffer size", "confidence": 0.71, "confidence_label": "validated", "answer_type": "answer"}
claude code → your expert
export ANTHROPIC_BASE_URL=https://serve.your-company.internalexport ANTHROPIC_API_KEY=<customer key>export ANTHROPIC_MODEL=module-x-expertclaude# also: POST /v1/chat/completions, /v1/messages# a refusal is HTTP 200, answer_type "refusal"
Plans

Start with one expert.
Grow into a team of them.

FreeSoloTeamBusinessEnterprise
Experts131020unlimited
Seats11520unlimited
Knowledge units2002,00020,000200,000unlimited
API keys1310unlimitedunlimited
Commercial use○●●●●
GPU Engine and Finetune Hub○●●●●
Serve and Central○○●●●
Audit logs○○○●●

Invite-only beta. Prices on request.

Private beta is open

Tell us which field your knowledge comes from. We send invite codes in small cohorts.

Request access

[email protected] · Support: [email protected]