You have a long PDF. A report, a contract, or a manual. You ask a model (the system that writes the next words) where a number lives, and it quotes a definition instead.
That fetch-then-answer setup is RAG (grab a passage, then answer from that passage). The usual grabber is a vector database (a store that finds text by numeric closeness). PageIndex (a tool that keeps the headings and walks them) is the other grabber.
On a PageIndex vs vector database choice, use PageIndex when the answer sits in one section of a long file with real headings. Use a vector database when you are searching a pile of short notes with no shared outline.
I ran both ideas on a one-page trap. Word overlap scored the glossary 12 and the real section 6. A walk that kept the headings returned 18.4 percent.
A long PDF is a stack of pages with a shape. The shape is the part most setups throw away.

Why does my RAG miss the right page in a long PDF

Your RAG misses the right page because it searched for similar words, and the glossary said those words more often than the results section did.
The model never opened the file. It only saw the passage your retriever handed over. If that passage is the definition of “operating margin,” the model will explain the ratio and never mention 18.4 percent.
This is the same split as a coding agent (a program that edits files, not a chat box that talks about code). The loop is not the model. If you want that split in plain words, read how a coding agent is the loop around the model.
Closeness in a number space. Two bits of text land near each other when an embedding (a list of numbers standing for that text) says they mean similar things. Near is not the same as “this is the line with the figure.”
What a vector database actually compares
A vector database compares number lists, then returns the chunks (sliced pieces of the file) whose lists sit nearest the question.
Here is the path, with each new piece named once.
- You cut the PDF into chunks. A chunk might be a paragraph, or a fixed number of tokens (small pieces of words the model is billed for).
- A smaller model turns each chunk into an embedding.
- The vector database stores those lists.
- Your question becomes a list too. The database returns the nearest lists.
Nothing in that path knows that “Glossary” is a definition and “Operating results” is a figure. If the definition repeats the phrase, it wins.
Pasting the whole PDF instead is the other failure. A context window (how much text the model can hold at once) is finite. A long filing does not fit. You also pay for every token you send. That bill is the same family of problem as a coding session that rereads the whole thread: why a full session gets expensive.
How PageIndex finds a section without embeddings
PageIndex finds a section by keeping the headings as a tree, then letting a model walk that tree the way you would use a table of contents.
I checked the installed code. The tree of a markdown file comes from the heading lines. That step does not call a model. The project says the same thing about PDFs: the outline comes from the layout, and a model is only used later to write a short note on each node (one heading plus the text under it) and to choose which node to open.

The chat step is different. Local mode asks a model to walk those nodes. I did not run that call. It needs an API key, and the quality depends on which model you can pay for. The project says a basic model is enough to build the tree, and you should spend on the model that searches it.
PageIndex vs vector database on one trap report
On this trap, word overlap picked the glossary and the heading walk picked 18.4 percent.
I wrote a tiny report called Northwind. The glossary says “operating margin” six times and states no percent. The operating results section says it twice and includes the line “The operating margin was 18.4 percent.”
I used PageIndex’s own extractor, extract_nodes_from_markdown, then build_tree_from_nodes. No key. It returned nodes 0001 through 0005, with 0002 Glossary and 0004 Operating results.
Then I scored the words operating and margin in the title string plus the section text. Glossary scored 12. Operating results scored 6. The heading line is stored inside the text and I also added the title, so the word Operating on the results node was counted twice. Count it once and the glossary still wins.
A second rule asked a question a person actually asks: which section has those words and also a digit. Only Operating results matched. The answer line was 18.4 percent.

| Word overlap | Heading walk | |
|---|---|---|
| What it looks at | How often the words repeat | Headings, then a number in the section |
| Winner on the trap | Glossary, score 12 | Operating results, 18.4 percent |
| Needs an API key | No | No, for this stand-in |
| What it is not | A full embedding model | PageIndex’s paid chat walk |
This stand-in is stricter than a real embedding model in one way and looser in another. A real embedding might still prefer the glossary, because the phrase is louder there. It might also be smart enough to notice “definition.” I did not claim either. I claimed the failure mode: similarity has no idea which heading is a figure.
The full PageIndex chat walk is the product version of the right-hand column. It can open a node, read it, and back up if the section is wrong. My digit rule cannot do that. It was enough to show why the tree exists.
How to build the tree without an API key
You can build the heading tree in a few minutes with the package and no key, as long as the file is markdown or you already have the headings.
pip install -U pageindexSave this as trap.md:
# Northwind Annual Report
## Glossary
The operating margin is a ratio. People also say operating margin when they mean profitability talk in general. This glossary repeats operating margin so a similarity search sees the phrase operating margin many times. Operating margin here is only a definition, not a result. No percent is stated in this operating margin glossary entry.
## Operating results
Revenue grew. The operating margin was 18.4 percent.Then run:
import re
from collections import Counter
from pageindex.page_index_md import (
extract_nodes_from_markdown,
extract_node_text_content,
build_tree_from_nodes,
)
md = open("trap.md", encoding="utf-8").read()
nodes, lines = extract_nodes_from_markdown(md)
nodes = extract_node_text_content(nodes, lines)
tree = build_tree_from_nodes(nodes)
def flat(ns):
for n in ns:
yield n
yield from flat(n.get("nodes") or [])
q = set("operating margin".split())
def words(n):
blob = (n["title"] + " " + n.get("text", "")).lower()
c = Counter(re.findall(r"[a-z0-9]+", blob))
return sum(c[w] for w in q)
for n in flat(tree):
print(n["node_id"], n["title"], "word_score", words(n))You should see Glossary above Operating results on word_score. Then look at which section text contains a digit. That is the whole test.
For a real PDF, the documented client is different. It wants OPENAI_API_KEY, and you pass the model names you can actually call:
import os
from pageindex import PageIndexClient
os.environ["OPENAI_API_KEY"] = "your-key"
client = PageIndexClient() # or pass index= and chat= model names you pay for
doc_id = client.submit_document("report.pdf")["doc_id"]
print(client.chat("What was the operating margin?", doc_id=doc_id))I stopped before that call. Do not paste a key into a chat log. Local mode, as shipped, reads text-based PDFs. It does not do OCR (reading words out of a picture of a page). A scan of a paper filing can come back with no headings at all.
Docs for the client live at the PageIndex developer docs. The code is in the PageIndex repo.
When not to use PageIndex
Do not use PageIndex for a large pile of short, messy notes. Use a vector database there.
PageIndex builds a tree per document. That is the right shape for a 200-page filing. It is the wrong shape for 50,000 support tickets, a chat archive, or a scrape of random web pages. Those have no shared table of contents. A cheap nearest-chunk lookup is what you want, and you can add a keyword pass beside it.

Also skip it when:
- The PDF is a scan and you have no text layer. Local mode will not invent headings from a photo.
- You need the answer in a few tenths of a second. A tree walk calls a model. A vector lookup does not.
- You have millions of queries a day on simple facts. The project itself shows query cost climbing with how much reasoning you buy. A vector lookup stays cheaper per question.
- The question is not in the file. No retriever fixes a model that has to guess.
The project reports 98.7% on FinanceBench (a public set of questions about financial filings) and labels a vector setup at 50% on that same chart. I did not re-run FinanceBench. Treat that gap as their claim about long financial documents, not as a promise for your ticket bot.

They also report that sending the whole PDF costs more as the file grows, about 2.1 times their retrieval at 52 pages and about 16.6 times at 420 pages, and that an 805-page file no longer fits the context window. I did not re-run that either. The shape matches the reason you retrieve at all: do not pay to reread every page for every question.

If you want both, do not mash them into one index. Search the pile with keywords or vectors to pick the document. Then walk that one document with a heading tree. The tree is a bad catalog. It is a good table of contents.
Try this on your own report
Run the script above on a real outline, not on Northwind, and see which heading word-overlap would have opened.
Take one PDF you already trust. Export the first few headings to markdown, or copy the table of contents into trap.md and put the real number under the right heading and a repeated phrase under Glossary. Run the script. If word overlap opens the wrong heading, a vector index of that file will be tempted by the same trap.
Then, if you have a key, run PageIndexClient on the PDF and ask the same question. Compare the section it cites with the section the word score picked. That is a ten-minute test. If both land on the right section, you may not need the tree for this file. If only the tree lands there, stop embedding that filing.
A decision tree is just a chain of ifs. Yours can be three lines: does the file have headings, is the corpus one document or many, do you need the section cited.
Common questions about PageIndex vs vector database
Can I do RAG without a vector database?
Yes. Keep the headings and walk them. That is what PageIndex does. You still need something that writes the answer. You do not need an embedding index for a single structured file.
Why does similarity search return the glossary instead of the number?
Because the glossary repeats the words in your question, and similarity is a count of closeness, not a check that the sentence contains the figure. On my trap the glossary scored 12 and the results section scored 6.
Is PageIndex worth it for a chatbot over lots of short docs?
No. Short docs and chat logs do not share an outline. A vector database, often plus a keyword pass, is the better default. Add a heading walk only after you have picked one long document.
Do I still need embeddings if my PDF has a table of contents?
Not for that PDF. The headings are already the index. Use embeddings when you are matching many files that do not share those headings, or when you need a fast shortlist before anyone walks a tree.




