Back to blog
#AI#RAG#JavaScript#embeddings#tutorial#beginners

Build a Document Q&A App — Your Second AI App (RAG in Plain JavaScript)

By DevKingOv8 min read
XLinkedIn

Your first AI app was a chat loop — a text box, a server route, an API, streaming back. Powerful, but it only knows what the model was trained on. The moment users ask "what does my document say?", you need your second app: document Q&A, the "chat with your data" pattern behind every enterprise AI product.

The technique has a name that scares beginners — retrieval-augmented generation — and a reality that doesn't: you find the relevant parts of the document and paste them into the prompt. That's the secret. Everything else is plumbing to do that finding well. This guide builds it in plain JavaScript, provider-neutral, as the natural sequel to building your first AI app — stage three of the learning roadmap, project two.

The problem RAG actually solves

Models have two hard limits: they've never seen your documents, and they can't reliably reason over text that isn't in the prompt. There are three ways around this — and the differences explain most RAG architecture decisions:

Stuff everything in the prompt Fine-tune the model RAG (retrieve + augment)
How it works Paste the whole doc into each request Train the model on your data Find relevant chunks, paste only those
Works for Docs under a few dozen pages Style/tone, narrow domains Real document collections
Cost per question High — you pay for all text, every time Training cost, then normal Low — only relevant chunks
Keeps data current Re-send everything Retrain on every change Re-index changed docs
Answers cite sources Yes, but buried No — it "just knows" Yes, by design
Beginner-feasible Yes No Yes — this guide

Fine-tuning is the one beginners assume they need; it's actually the rarest fit. RAG wins for document Q&A because knowledge updates by re-indexing a file, not retraining a model — and because every answer can point at the exact passage it came from.

The pipeline, in one picture

[ Document ] → split into chunks → embed each chunk → store vectors
                                                        │
[ Question ] → embed the question ──────────────────────┤
                                                        ↓
                              find the most similar chunks (vector search)
                                                        ↓
                        [ prompt: "Answer using ONLY these excerpts" ] → AI → cited answer

Five steps: chunk → embed → store → retrieve → assemble. Let's build each in plain JS.

Step 1 — Chunking: split documents sensibly

Feed a model a 100-page PDF in one prompt and it drowns; feed it random 5-word fragments and it can't see context. Chunking is the middle path: split documents into pieces small enough to retrieve individually but large enough to carry meaning.

function chunk(text, size = 800, overlap = 100) {
  const chunks = [];
  let start = 0;
  while (start < text.length) {
    chunks.push(text.slice(start, start + size));
    start += size - overlap; // overlap so boundaries don't cut ideas in half
  }
  return chunks;
}

Two knobs: size (hundreds of words, roughly a paragraph or two) and overlap (so a sentence split at a boundary appears whole in at least one chunk). For your first version, split on paragraph boundaries first and merge until near the size limit — structure-aware chunking beats character math.

Keep each chunk's provenance: { docId, text, page: 3, heading: "Pricing" }. Without provenance you can't cite, and without citations users can't trust — more on that at the end.

Step 2 — Embeddings: meaning as coordinates

An embedding turns a piece of text into a vector — a list of a few hundred numbers — such that texts with similar meaning land near each other in that space. "How do I cancel?" sits close to "cancel my subscription," even with zero words in common. That's the entire trick, and you use it via an API call:

// Provider-neutral shape — check your provider's current embeddings docs.
async function embed(texts) {
  const res = await fetch("https://YOUR-AI-PROVIDER/v1/embeddings", {
    method: "POST",
    headers: {
      "Content-Type": "application/json",
      Authorization: `Bearer ${process.env.AI_API_KEY}`, // server-side only, always
    },
    body: JSON.stringify({ model: "an-embedding-model", input: texts }),
  });
  const { data } = await res.json();
  return data.map((d) => d.embedding); // array of number-arrays
}

Same security rule as your first app: this route runs on the server, the key never ships to the browser.

Step 3 — Storage: start embarrassingly simple

Beginners burn weeks choosing vector databases. Don't. Your first version needs a cosine-similarity loop over an array:

// Good enough for thousands of chunks. Honest about its limits.
const index = []; // { vector: number[], text, docId, page }

function cosine(a, b) {
  let dot = 0, na = 0, nb = 0;
  for (let i = 0; i < a.length; i++) {
    dot += a[i] * b[i];
    na += a[i] * a[i];
    nb += b[i] * b[i];
  }
  return dot / (Math.sqrt(na) * Math.sqrt(nb));
}

function search(queryVec, k = 4) {
  return index
    .map((c) => ({ ...c, score: cosine(queryVec, c.vector) }))
    .sort((x, y) => y.score - x.score)
    .slice(0, k);
}

In-memory (or a JSON file) carries a personal project a long way. When it doesn't — many documents, many users, filters by tenant — graduate to a real vector store; if you're already running Postgres, its vector extension handles this natively, and your schema comes along for the ride.

Step 4 — Retrieval: embed the question, find its neighbors

At question time, embed the question with the same model (this matters — different models produce incompatible spaces), run the similarity search, and take the top few chunks:

const [qVec] = await embed([question]);
const hits = search(qVec, 4); // the chunks most similar in meaning

Resist tuning k early. Retrieval quality is mostly decided by chunking and, later, by metadata filters ("only search the 2026 handbook") — not by grabbing more chunks.

Step 5 — Assembly: the prompt makes or breaks it

Now the actual augmentation — the retrieved context meets the model:

const context = hits
  .map((h, i) => `[${i + 1}] (${h.docId} p.${h.page}) ${h.text}`)
  .join("\n\n");

const messages = [
  {
    role: "system",
    content:
      "Answer ONLY from the excerpts. Cite excerpt numbers like [2]. " +
      "If the excerpts don't contain the answer, say so plainly — never guess.",
  },
  { role: "user", content: `Excerpts:\n${context}\n\nQuestion: ${question}` },
];

Then stream it exactly like the chat app you already built. Three things in that system prompt are doing heavy lifting, and each is a prompt-engineering pattern in miniature: "ONLY" scopes the model to your data, citations make answers verifiable, and permission to say "not in the docs" is your single best defense against hallucinated answers. Users trust "I don't know" far more than a confident wrong answer.

Where RAG fails (and your fixes)

  • Answers miss the point — usually chunking, not retrieval: chunks too small (context shredded) or too big (noise retrieved). Fix boundaries before touching anything else.
  • Right chunk, wrong doc — a stale index. Re-index on document change; version your chunks by document hash so you know when to.
  • Confident answers from nowhere — the system prompt lost its guardrails, or retrieval returned nothing relevant and the model filled the void. Log retrieval scores; when top scores are low, short-circuit to "couldn't find this in your documents."
  • It works in the demo, dies with real files — real documents have tables, scans, and weird encodings. Extraction is its own step; test early with the ugliest file a user has ever sent you.

Ship the simple version, watch real questions, and fix what actually fails — the same review-the-diff discipline applied to a pipeline instead of code.

FAQ

Do I need machine learning to build a RAG app?

No. Every ML-shaped step — embeddings, similarity — happens behind an API or a 10-line loop. What you need is solid JavaScript: string handling, fetch, and state. RAG is an architecture pattern, not a model you train.

RAG vs fine-tuning — which should a beginner learn?

RAG, decisively. It solves the common case (making specific documents answerable), updates by re-indexing instead of retraining, produces citable answers, and needs zero ML infrastructure. Fine-tuning is a later, specialized tool for changing model behavior or style at scale — learn it when a problem actually demands it.

How many chunks should I retrieve per question?

Start with three to five reasonably-sized chunks and tune only if evidence says so. Retrieval quality is dominated by chunking quality and metadata filtering, not by count. Retrieving "more" mostly adds noise and cost.

Can I run document Q&A fully offline?

Yes — open-weight embedding and chat models run locally, and the pipeline above stays identical (swap the API host). The trade-offs are setup effort and hardware. For learning and shipping fast, hosted APIs are the simpler start; local is a great second milestone.

What should my third AI app be?

A mini agent — a loop where the model calls your functions and decides the next step. It's the natural third project in the roadmap, and the point where an architecture review from a mentor pays for itself. The DevKingOv courses walk the full sequence — chat, RAG, agents — project by project.

Prefer watching?

Every post here is a lesson in a free video course — follow along on YouTube and track your progress on the portal.

Keep reading

Want help applying this? Book a 1-on-1 with a consultant.

Find a Consultant