AI Knowledge Assistant

Your company knowledge, answered with citations — a production RAG pipeline with permissions, evals, and monitoring.

Price
From $25K
Timeline
4–8 weeks
Terms
Fixed scope
The engagement

We build retrieval that cites its sources, respects permissions, and improves over time.

Ingestion, embeddings, hybrid retrieval, reranking, a citation format, a permissions model, a golden dataset and eval harness, a feedback loop, and monitoring — across four to eight weeks.

The opposite of chat-with-docs: a knowledge system tuned against a golden set you own, so answers are traceable, scoped, and get better instead of silently rotting.

Phase by phase
Week 101

Sources + golden set

We map your sources and build a golden dataset — the questions that must be answered right — before writing a line of retrieval.

Weeks 2–402

Ingestion + retrieval

Ingestion, chunking, embeddings, hybrid retrieval, and reranking — tuned against the golden set, not vibes.

Weeks 5–603

Citations + permissions

A citation format, a permissions model, and the guardrails that keep retrieval honest and scoped.

Weeks 7–804

Evals + monitoring

The eval harness, a feedback loop, and monitoring — so retrieval improves over time instead of rotting.

The work in context

An answer with a source. A handoff with context.

Explore an illustrative workflow. This is a design example, not a live agent, client deployment or measured result.

Human ownership
Human reviewReview requested
The decision belongs to a person

“Is the repair covered?”

Coverage needs the team’s judgement.
The caller asked for a callback.

Included in the handoff

Request · Call context · Callback details

Service managerAwaiting ownership
Illustrative interface · no real customer data

A person takes over.

The request travels with its context. Ownership is explicit.

Approved knowledge
Team knowledgeApproved sources

What happens if a caller asks about coverage?

Answer

Hand the request to the service manager. Do not make a coverage promise.

Illustrative answer and source · not live retrieval

An answer with its source.

The reference belongs beside the answer, not behind a claim of confidence.

Scope

What the engagement covers

Ingestion + chunking01
  • Source connectors
  • Structure-aware chunking
  • Incremental updates

Your knowledge stays fresh and chunked so retrieval actually finds the right passage.

Embeddings + retrieval02
  • Embedding model choice
  • Hybrid vector + keyword
  • Metadata filters

Retrieval finds the right passage even when the wording doesn't match.

Reranking03
  • Cross-encoder rerank
  • Relevance tuning
  • Top-k discipline

The best passages rise to the top, so the model answers from signal, not noise.

Citations04
  • Inline source links
  • Passage-level provenance
  • 'I don't know' honesty

Every answer is traceable to a source — and it says so when it can't find one.

Permissions05
  • Row / doc-level ACL
  • Per-user retrieval
  • No cross-tenant leakage

Users only retrieve what they're allowed to see — no leakage across tenants.

Evals + feedback06
  • Golden dataset
  • Recall + precision scoring
  • Thumbs → training

Retrieval is measured and improves over time, instead of silently degrading.

What you keep

What you receive

Tangible artifacts you keep, whether or not you continue past this engagement.

Deliverables · 11 included
  1. 01Source ingestion pipeline
  2. 02Embedding workflow
  3. 03Hybrid retrieval
  4. 04Reranking
  5. 05Citation format
  6. 06Permissions model
  7. 07Golden dataset
  8. 08Eval harness
  9. 09Feedback loop
  10. 10Monitoring
  11. 11Runbook

A retrieval system, delivered cited and evaluated — not chat-with-docs.

The pipeline that turns your knowledge into cited answers, the permissions that keep it honest, and the evals that keep it improving. A sample is shown; yours runs on your corpus.

A cited answer

The model answers from your sources, with inline provenance — and refuses when it can't.

retrieval · citedInteractive specimen
RequestStep 01

What's our refund window for enterprise?

The run record keeps the request beside the response.

Illustrative workflow. Select a stage to inspect it; this is not a client run.

Retrieval evals

Recall and precision against a golden dataset, tracked over time.

evals/retrieval.jsonEvaluation template
Golden setEvidence before a verdict
Not run
Recall @5

A representative input, expected outcome and execution trace are required before this dimension receives a result.

No grade or performance claim is shown before the buyer's evaluation.

The pipeline

Ingestion to answer, every stage tuned against the golden set.

pipeline · ragIllustrative states
Selected caseingestion + chunking

structure-aware

Input
Representative case
Evidence
Trace and expected outcome

The states demonstrate the review UI. They are not evaluation results.

What we measure

Agree the baseline, source and acceptance criteria before release.

Measurement contractNo sample outcomes
Selected measureAnswer accuracy
Baseline
Measured with the buyer
Source
Named system of record
Acceptance
Agreed before release

Targets and observed results belong to the qualified workflow, not a sample dashboard.

Illustrative product scenes · never client results

Plus: A golden dataset you own · a monitoring dashboard · 30 days of async Q&A.

Fit

Built for

Head of Support

Replacing keyword search over docs

Answers with citations your team trusts — not a search box that returns ten links.

AI PM

Grounding an agent in proprietary knowledge

The retrieval layer a custom agent needs to answer from your data reliably.

Data lead

Retrieving over sensitive knowledge

Permissioned, cited retrieval that respects who's allowed to see what.

FAQ

Questions, answered

Common questions

Start here

Start with one useful decision.

Bring the workflow and the person responsible. We scope the exact number with you, and sign a mutual NDA before any code or data is shared.

From $25K · 4–8 weeks