Teach a model
what you know.
Your knowledge as a searchable corpus, a stateless model inside an agent loop, and two kinds of training that make both sides better. Called like any language model.
From documents
to vectors.
The knowledge stays text. To find the right piece later, every piece is cut to a size one embedding can represent and turned into a vector.
- importPDF, Word, Markdown, text · CSV, Excel
- sections~2,400 chars, split at paragraphs, 200 chars overlap
- unitsAssist extracts typed units; table cells become facts without a model
- chunksunits over 4,000 chars split to ~1,600: paragraph → sentence → hard cut
- embedQwen3-Embedding-0.6B, 1024 dims
- indexone float32 matrix in memory + a BM25 index
| why chunk | one vector for a long text is an average of everything in it; a manual of 87 KB embedded as one unit matched nothing well |
|---|---|
| what is embedded | title + content; a fact as its fact line; a case as its situation and criteria, the problem rather than the write-up |
| search speed | one matrix-vector product: about 10 ms at 50,000 units |
| calibration | raw cosine is mapped per agent so the best relevant/irrelevant separator lands at 0.30 and the median relevant hit at 0.60, whatever the embedding model |
Six unit types: fact concept principle skill process case, linked by 14 edge types such as causes, requires, replaces or contradiction.
The model can't read everything.
So it gets what fits the question.
A context window holds a few thousand words of knowledge, a corpus holds millions. Retrieval picks the handful of units that answer this question, so the model works from your text instead of from what it half remembers.
- understandquestion type: lookup · explain · diagnose · advise · instruct · review
- densemeaning: cosine over the vectorsBM25exact words, codes, numbers
- RRFΣ 1/(60 + rank) merges both rankings
- rerankoptional cross-encoder over 30 candidates
- top 5score ≥ 0.25, until the knowledge budget is full
- graph≤ 2 hops along the edges the question type needs
top_score < 0.30 → refusal
If nothing scores high enough and there is no core knowledge or tool to answer from, the model is not even asked.
core: always in the prompt
The few rules and concepts every answer needs skip retrieval and sit in the system prompt, inside their own token budget.
diagnose → causes, mitigates, requires
A diagnosis pulls in causes and fixes, an instruction its prerequisites. An outdated unit whose replacement was found drops to the end.
The model is stateless.
The request is not.
A language model remembers nothing between calls. Nemonix gives one request a state of its own: the retrieved knowledge, the tool results gathered so far and a round counter. When the answer is out, that state is gone and the next request starts clean.
| tool loop | up to 4 rounds (1–20); when they are used up, one last call without tools forces the answer. The loop is bounded by Nemonix, not by the model |
|---|---|
| hidden tools | MCP servers you attach to the expert, such as analysis or lookup tools. Callers only ask questions; they never see, provide or reach these tools |
| when tools run | per tool: freely in the loop, only for certain question types, or always once before answering |
| tool results | enter the context as fenced data with their own token budget, never as instructions |
| caller history | the API is stateless: the caller sends the conversation with each request |
The units retrieved for this question, under “## VERIFIED KNOWLEDGE”, “## EXPERIENCE CASES” and “## RELATED KNOWLEDGE”. Hits are added in score order until this budget is full.
After the model answers, a second call checks it claim by claim against the knowledge and removes what isn't covered; sources are then attached without a model.
Teach the search
your vocabulary.
A general embedding model doesn't know that two of your terms mean the same thing. From 100 units on, it can be fine-tuned on your own corpus in the Finetune Hub, on your own GPU.
[question, unit, hard negative]
Questions come from a generator, from people and from your test set. The hard negative is the nearest unit that has no edge to the right one.
MultipleNegativesRankingLoss
Contrastive: the right unit has to beat its hard negative and every other unit in the batch. One epoch, batch 16, 20 % held out.
recall@5 ↑ and nDCG@5 ≥
Otherwise the old model stays. Adopting re-embeds the corpus and recalibrates the scores.
* Internal measurement on one test corpus, July 2026. Your numbers depend on your knowledge.
Train how it uses knowledge,
not the knowledge itself.
A full fine-tune rewrites every weight, forgets more and leaves one large model per customer. A LoRA adapter trains about 2 % of the parameters on a frozen base; with QLoRA that base is held in 4 bit. It learns less, and that is fine: the facts it should apply arrive in the context with every question, through retrieval and core knowledge. Trained in the Finetune Hub, on your own GPU.
| method | LoRA, or QLoRA with the base in 4-bit NF4; adapters on every linear layer, rank 16 |
|---|---|
| base model | your choice: any open model you can run |
| what it learns | to read the given knowledge, copy values exactly, reason the way your expert does, and refuse what the context doesn't cover |
| counterfactuals | examples change a known value in the context, so the model learns to trust the text over its memory |
| cost | an adapter is tens to low hundreds of MB over one shared base, instead of a full model copy per customer (~54 GB at 27B) |
| memory | estimated ~25 GB to train a 27B base, ~10 GB for 8B |
An AI that helps you
build the expert.
Assist is an agent of its own inside the Studio. It reads what is already there, proposes new knowledge, and tests the expert on a sample question, while your expert stays the one who accepts.
| proposes | typed units with synonyms and a core suggestion, topics, edits and archivals, and pipeline changes |
|---|---|
| case interview | asks for a real decision in five parts, the way experts recall critical decisions |
| where it lands | the Inbox: written at once, marked validated: false, accepted by a person |
| model | your own key, or managed: Claude Haiku from Solo, Sonnet from Team, Opus from Business |
Outside agents can propose knowledge too, over MCP. They land in the same Inbox and can't validate anything themselves.
The Studio builds.
The modules run it.
The desktop app: knowledge, pipeline, Playground, test questions, releases.
Always-on runtime for released versions. A version freezes knowledge, embeddings, model and pipeline.
Local embeddings on your own GPU, from 6 GB of VRAM.
A private workspace server for your team, on any host in your network.
Fine-tune the retriever and the LLM, on your own hardware.
Called like any model.
Serve speaks the OpenAI and Anthropic formats. Claude Code or OpenCode can use an expert as their model; their own tools are chosen by Serve but run on the caller's side, never inside the expert.
Start with one expert.
Grow into a team of them.
| Free | Solo | Team | Business | Enterprise | |
|---|---|---|---|---|---|
| Experts | 1 | 3 | 10 | 20 | unlimited |
| Seats | 1 | 1 | 5 | 20 | unlimited |
| Knowledge units | 200 | 2,000 | 20,000 | 200,000 | unlimited |
| API keys | 1 | 3 | 10 | unlimited | unlimited |
| Commercial use | ○ | ● | ● | ● | ● |
| GPU Engine and Finetune Hub | ○ | ● | ● | ● | ● |
| Serve and Central | ○ | ○ | ● | ● | ● |
| Audit logs | ○ | ○ | ○ | ● | ● |
Invite-only beta. Prices on request.
Private beta is open
Tell us which field your knowledge comes from. We send invite codes in small cohorts.
Request access[email protected] · Support: [email protected]