
Giving a Model Your Own Documents
A model that knows your product, your policies and your handbook is more useful than one answering from general training. There are two ways to get there, and choosing between them is most of the work. This post covers both, plus the chunking, citation and testing that decide whether the result is trustworthy.
The two approaches
The first is to put the material in the prompt. If the relevant document is a handful of pages, paste it in, ask the question, and you are done. No infrastructure, no index, no surprises about what the model did or did not see.
The second is retrieval: keep the documents in a searchable store, find the passages that relate to the question, and put only those in the prompt. This is what people mean by RAG. You need it once the material is larger than the context window, changes often, or has to respect who is allowed to see what.
Start with the first. Plenty of internal tools that were built as retrieval systems would have worked better as a prompt containing one well-maintained document.
When everything fits in the context
Long context windows make the simple approach viable for surprisingly large material. Two things still matter. Cost and latency scale with what you send, so sending a hundred pages to answer a one-line question is expensive per query even when it works. And attention is not uniform: models are reliably better with material at the start and end of a long context than with something buried in the middle. If you know which section matters, put it near the question.
Prompt caching, offered by several providers, changes the economics when the same large document is sent repeatedly. The stable prefix is cached and charged at a lower rate, which makes "paste the handbook every time" affordable in a way it was not before.
Retrieval, in plain terms
Retrieval has four steps. Split the documents into chunks. Turn each chunk into a vector with an embedding model, which places similar text near each other in a mathematical space. Store those vectors. At query time, embed the question, find the nearest chunks, and put them in the prompt.
Two practical notes. Vector search finds text that means something similar, which is what you want for questions phrased differently from the source. It is weaker at exact matches: product codes, error numbers, names. Most serious systems combine it with a keyword search and merge the results, which is worth doing from the start rather than after the first complaint about a part number.
And you do not always need a specialist database. Postgres with pgvector, or the vector support in the search engine you already run, will carry a substantial corpus. Add a dedicated store when you have a measured reason.
Chunking that respects the document
Chunking by character count cuts sentences in half and separates a heading from the paragraph it introduces. Split on structure instead: sections, headings, paragraphs. Keep chunks in the region of a few hundred words, with a small overlap so a sentence spanning a boundary survives.
The single most effective trick is to prepend context to each chunk before embedding it: the document title, the section heading, sometimes a one-line summary of what the document is. A chunk that begins "Refund policy, section 3: exceptions" retrieves far better than one that begins "This does not apply when".
Store the source with each chunk: document name, section, a link, and a last-updated date. You need them for citations and for the moment someone asks where an answer came from.
Make the answer point at the source
An answer without a source is a claim you have to verify by hand, which removes most of the time saving. Ask for the citation in the prompt, and make the refusal path explicit:
Answer using only the passages below.
Quote the sentence you relied on, and name its source.
If the passages do not contain the answer, say so and stop.
That last instruction is what stops a confident answer assembled from nothing. It also gives you a measurable signal: a rising share of "not covered" responses usually means the retrieval step is failing, not that the model got worse.
Keeping it current
An index is a copy, and copies go stale. Decide how a change to a document reaches the index: a nightly rebuild is fine for a handbook, an event on save is better for anything people edit daily. Keep the last-updated date with each chunk and show it in the answer, so a reader can see they are looking at something from last March.
Deletions matter more than they seem. A document that was withdrawn but is still in the index will keep being quoted back at people as current policy.
Two failures, tested separately
When an answer is wrong, there are two possible causes, and they need different fixes: the right passage was never retrieved, or it was retrieved and the model used it badly. Test them separately.
Build a set of real questions with the passage that should answer each. Measure retrieval first: for each question, is the correct chunk in the top results? That number is fixed by chunking, embeddings and hybrid search. Only then look at the answers, where the fixes are prompt-level: citation requirements, output format, the refusal instruction.
Twenty to fifty questions taken from what people actually ask is enough to see both numbers move. Rerun the set after every change to chunking or prompt, because improvements in one direction routinely break the other.
The parts that quietly go wrong
- Extraction. Text pulled out of a PDF arrives with broken line breaks, headers repeated on every page, and tables flattened into unreadable rows. Check what your extractor produces before blaming the model.
- Scans. An image-only PDF contains no text at all until it goes through OCR.
- Permissions. If your documents have access rules, the retrieval step has to apply them per user. A shared index that ignores permissions will happily quote the salary review to the wrong person.
- Confidential material. Whatever you send leaves your systems unless you are running the model yourself. Check the terms and your own obligations before indexing personnel or customer files.
Get those right and the rest is tuning. Get them wrong and no amount of prompt work will rescue the result.
Comments
No comments yet. Be the first to share your thoughts.


