Build a Document Q&A App — Your Second AI App (RAG in Plain JavaScript)
Your first AI app was a chat loop — a text box, a server route, an API, streaming back. Powerful, but it only knows what the model was trained on. The moment users ask "what does my document say?", you need your second app: document Q&A, the "chat with your data" pattern behind every enterprise AI product.
The technique has a name that scares beginners — retrieval-augmented generation — and a reality that doesn't: you find the relevant parts of the document and paste them into the prompt. That's the secret. Everything else is plumbing to do that finding well. This guide builds it in plain JavaScript, provider-neutral, as the natural sequel to building your first AI app — stage three of the learning roadmap, project two.
The problem RAG actually solves
Models have two hard limits: they've never seen your documents, and they can't reliably reason over text that isn't in the prompt. There are three ways around this — and the differences explain most RAG architecture decisions:
| Stuff everything in the prompt | Fine-tune the model | RAG (retrieve + augment) | |
|---|---|---|---|
| How it works | Paste the whole doc into each request | Train the model on your data | Find relevant chunks, paste only those |
| Works for | Docs under a few dozen pages | Style/tone, narrow domains | Real document collections |
| Cost per question | High — you pay for all text, every time | Training cost, then normal | Low — only relevant chunks |
| Keeps data current | Re-send everything | Retrain on every change | Re-index changed docs |
| Answers cite sources | Yes, but buried | No — it "just knows" | Yes, by design |
| Beginner-feasible | Yes | No | Yes — this guide |
Fine-tuning is the one beginners assume they need; it's actually the rarest fit. RAG wins for document Q&A because knowledge updates by re-indexing a file, not retraining a model — and because every answer can point at the exact passage it came from.
The pipeline, in one picture
[ Document ] → split into chunks → embed each chunk → store vectors
│
[ Question ] → embed the question ──────────────────────┤
↓
find the most similar chunks (vector search)
↓
[ prompt: "Answer using ONLY these excerpts" ] → AI → cited answer
Five steps: chunk → embed → store → retrieve → assemble. Let's build each in plain JS.
Step 1 — Chunking: split documents sensibly
Feed a model a 100-page PDF in one prompt and it drowns; feed it random 5-word fragments and it can't see context. Chunking is the middle path: split documents into pieces small enough to retrieve individually but large enough to carry meaning.
function chunk(text, size = 800, overlap = 100) {
const chunks = [];
let start = 0;
while (start < text.length) {
chunks.push(text.slice(start, start + size));
start += size - overlap; // overlap so boundaries don't cut ideas in half
}
return chunks;
}
Two knobs: size (hundreds of words, roughly a paragraph or two) and overlap (so a sentence split at a boundary appears whole in at least one chunk). For your first version, split on paragraph boundaries first and merge until near the size limit — structure-aware chunking beats character math.
Keep each chunk's provenance: { docId, text, page: 3, heading: "Pricing" }. Without provenance you can't cite, and without citations users can't trust — more on that at the end.
Step 2 — Embeddings: meaning as coordinates
An embedding turns a piece of text into a vector — a list of a few hundred numbers — such that texts with similar meaning land near each other in that space. "How do I cancel?" sits close to "cancel my subscription," even with zero words in common. That's the entire trick, and you use it via an API call:
// Provider-neutral shape — check your provider's current embeddings docs.
async function embed(texts) {
const res = await fetch("https://YOUR-AI-PROVIDER/v1/embeddings", {
method: "POST",
headers: {
"Content-Type": "application/json",
Authorization: `Bearer ${process.env.AI_API_KEY}`, // server-side only, always
},
body: JSON.stringify({ model: "an-embedding-model", input: texts }),
});
const { data } = await res.json();
return data.map((d) => d.embedding); // array of number-arrays
}
Same security rule as your first app: this route runs on the server, the key never ships to the browser.
Step 3 — Storage: start embarrassingly simple
Beginners burn weeks choosing vector databases. Don't. Your first version needs a cosine-similarity loop over an array:
// Good enough for thousands of chunks. Honest about its limits.
const index = []; // { vector: number[], text, docId, page }
function cosine(a, b) {
let dot = 0, na = 0, nb = 0;
for (let i = 0; i < a.length; i++) {
dot += a[i] * b[i];
na += a[i] * a[i];
nb += b[i] * b[i];
}
return dot / (Math.sqrt(na) * Math.sqrt(nb));
}
function search(queryVec, k = 4) {
return index
.map((c) => ({ ...c, score: cosine(queryVec, c.vector) }))
.sort((x, y) => y.score - x.score)
.slice(0, k);
}
In-memory (or a JSON file) carries a personal project a long way. When it doesn't — many documents, many users, filters by tenant — graduate to a real vector store; if you're already running Postgres, its vector extension handles this natively, and your schema comes along for the ride.
Step 4 — Retrieval: embed the question, find its neighbors
At question time, embed the question with the same model (this matters — different models produce incompatible spaces), run the similarity search, and take the top few chunks:
const [qVec] = await embed([question]);
const hits = search(qVec, 4); // the chunks most similar in meaning
Resist tuning k early. Retrieval quality is mostly decided by chunking and, later, by metadata filters ("only search the 2026 handbook") — not by grabbing more chunks.
Step 5 — Assembly: the prompt makes or breaks it
Now the actual augmentation — the retrieved context meets the model:
const context = hits
.map((h, i) => `[${i + 1}] (${h.docId} p.${h.page}) ${h.text}`)
.join("\n\n");
const messages = [
{
role: "system",
content:
"Answer ONLY from the excerpts. Cite excerpt numbers like [2]. " +
"If the excerpts don't contain the answer, say so plainly — never guess.",
},
{ role: "user", content: `Excerpts:\n${context}\n\nQuestion: ${question}` },
];
Then stream it exactly like the chat app you already built. Three things in that system prompt are doing heavy lifting, and each is a prompt-engineering pattern in miniature: "ONLY" scopes the model to your data, citations make answers verifiable, and permission to say "not in the docs" is your single best defense against hallucinated answers. Users trust "I don't know" far more than a confident wrong answer.
Where RAG fails (and your fixes)
- Answers miss the point — usually chunking, not retrieval: chunks too small (context shredded) or too big (noise retrieved). Fix boundaries before touching anything else.
- Right chunk, wrong doc — a stale index. Re-index on document change; version your chunks by document hash so you know when to.
- Confident answers from nowhere — the system prompt lost its guardrails, or retrieval returned nothing relevant and the model filled the void. Log retrieval scores; when top scores are low, short-circuit to "couldn't find this in your documents."
- It works in the demo, dies with real files — real documents have tables, scans, and weird encodings. Extraction is its own step; test early with the ugliest file a user has ever sent you.
Ship the simple version, watch real questions, and fix what actually fails — the same review-the-diff discipline applied to a pipeline instead of code.
FAQ
Do I need machine learning to build a RAG app?
No. Every ML-shaped step — embeddings, similarity — happens behind an API or a 10-line loop. What you need is solid JavaScript: string handling, fetch, and state. RAG is an architecture pattern, not a model you train.
RAG vs fine-tuning — which should a beginner learn?
RAG, decisively. It solves the common case (making specific documents answerable), updates by re-indexing instead of retraining, produces citable answers, and needs zero ML infrastructure. Fine-tuning is a later, specialized tool for changing model behavior or style at scale — learn it when a problem actually demands it.
How many chunks should I retrieve per question?
Start with three to five reasonably-sized chunks and tune only if evidence says so. Retrieval quality is dominated by chunking quality and metadata filtering, not by count. Retrieving "more" mostly adds noise and cost.
Can I run document Q&A fully offline?
Yes — open-weight embedding and chat models run locally, and the pipeline above stays identical (swap the API host). The trade-offs are setup effort and hardware. For learning and shipping fast, hosted APIs are the simpler start; local is a great second milestone.
What should my third AI app be?
A mini agent — a loop where the model calls your functions and decides the next step. It's the natural third project in the roadmap, and the point where an architecture review from a mentor pays for itself. The DevKingOv courses walk the full sequence — chat, RAG, agents — project by project.
Prefer watching?
Every post here is a lesson in a free video course — follow along on YouTube and track your progress on the portal.
Keep reading
Build Your First AI App with JavaScript — A Beginner's Guide
The complete path to your first deployed AI-powered app: how AI APIs actually work, the safe server-route pattern, streaming responses, and which first project to pick. No machine learning required — just JavaScript.
AI Pair Programming for Beginners — How to Review AI-Generated Code
The skill that decides whether AI makes you faster or buries you in bugs: a beginner's guide to reviewing AI-generated code, prompting for reviewable output, and knowing when to take over.
How to Learn AI-Assisted Coding in 2026 — A Practical Roadmap
The realistic path from zero to shipping AI-powered apps: fundamentals first, then AI pair programming, then building with AI APIs. A stage-by-stage roadmap with a 12-week plan, tool choices, and the mistakes that stall most beginners.
Want help applying this? Book a 1-on-1 with a consultant.
Find a Consultant