Kestrel Leasing just launched a new lease tier, and support call volume tripled overnight. Almost every call has the same shape: a driver who signed a 40-page lease months ago, calling to ask “what’s my mileage cap?” or “can I end this early without a fee?” Hold times hit 20–40 minutes for a question with a one-line answer — if you know where to look in the contract.
This isn’t just a car-leasing problem. It’s the same shape wherever a customer already holds the document that answers their own question: an insurance policyholder asking “am I covered for this?”, a telecom customer asking about early-exit fees, anyone Ctrl-F-ing their way through a contract they signed months ago while a support queue backs up behind them. If any of that sounds like your own team, keep reading — the fix generalizes.
First, what “RAG” actually means
Kestrel’s fix is built on Amazon Bedrock Knowledge Bases, using a pattern called RAG — Retrieval-Augmented Generation. Stripped of the acronym: it’s an open-book exam, not a closed-book one. A model with no RAG answers from memory, the way a student tries to recall a 40-page contract they read once, months ago. RAG hands the model the one relevant page, open, right before it has to answer. That’s the whole idea — look it up, then answer, never answer from memory alone.
For Kestrel, that means the webchat never asks the model “what does a typical lease say about early termination?” from memory. It finds the one clause in that specific customer’s lease first, hands the model just that clause, and only then asks for an answer.
The interesting part isn’t that they used RAG — plenty of teams reach for it. It’s the decision they had to get right underneath it: how do you split a 40-page legal document into pieces a model can actually search?

Why you can’t just hand the model the whole contract
Foundation models today can technically fit a 40-page document in context. That’s not the problem. The problem is retrieval precision: ask “can I end my lease early?” and a whole-document match returns the entire contract as “relevant” — which isn’t an answer a webchat customer will wait through. The model needs to point at the early-termination clause specifically, not summarize the whole lease back at them.
Kestrel tried the obvious shortcut first — a simple FAQ bot with ten canned answers. It broke on the first real question, because “what’s my mileage cap” depends on which lease that customer actually signed. A keyword match can’t tell two customers’ contracts apart. That failure is what actually justified reaching for a model at all, rather than assuming AI was the answer from the start.
That’s where chunking comes in — and it’s easy to confuse with two other steps that sound similar but do different jobs:
- Tokenization — the model breaking text into word-pieces internally, on every request. Happens automatically; not something you configure.
- Embedding — turning a chunk of text into a vector so “mileage cap” and “kilometer limit per year” land near each other in meaning. This happens after chunking, not instead of it.
- Chunking — the actual split. Before anything else happens, the document gets cut into smaller, retrievable pieces. This is the step Kestrel actually had to design.
Chunking isn’t one decision — it’s a family of them
Kestrel’s real document set made this concrete: newer leases follow a clean, numbered template (sections → subsections → clauses). Older leases were scanned and OCR’d from paper, with unreliable section numbering. One chunking strategy doesn’t fit both.
Hierarchical chunking splits along the document’s own structure. It creates a two-level split — a precise small chunk (one clause) for matching, and a larger parent chunk (the full section) for context — so the model gets both accuracy and enough surrounding detail to answer well. This is the right fit for Kestrel’s newer, well-structured leases.
Semantic chunking splits by meaning instead of structure — grouping sentences that are actually about the same idea, even when the document’s own numbering can’t be trusted. This is the right fit for the older, scanned leases. It’s also more expensive to run: AWS’s own documentation is explicit that semantic chunking carries additional cost because it uses a foundation model at ingestion time, unlike hierarchical or fixed-size chunking.
Kestrel’s actual decision: two chunking strategies, routed by document type — hierarchical for the modern template, semantic for the older scanned contracts — both landing in the same vector store, so retrieval never needs to know which path a chunk came from.
The part most write-ups skip: how a lease actually gets in
It’s easy to draw a diagram that jumps straight from “customer asks a question” to “system searches S3” — but that skips two real questions any developer building this would actually have to answer.

First: how does a lease reach S3 at all? Two real sources feed the same bucket here — Legal/Ops uploads a newly signed lease, or an older paper lease comes out of a digitization backlog (scanned and OCR’d). Both land as a PDF in the same S3 prefix; the split by era happens later, at chunking.
Second: what actually triggers ingestion once a file is sitting in S3? This is the part worth being precise about: StartIngestionJob is not something Bedrock Knowledge Bases calls automatically. AWS’s own documentation is explicit that it’s a manually-invoked API call — no built-in trigger fires it for you. The standard way teams close that gap, and the pattern AWS itself documents in an official ML blog post, is an S3 Event Notification (fires the moment a file lands, for near-real-time pickup) or an EventBridge Scheduler rule (a nightly catch-up sync, or a run kicked off whenever a policy changes) — either one invoking a small Lambda function that calls StartIngestionJob on your behalf. Real detail worth knowing if you’re building this yourself: that API has a rate limit of 0.1 requests per second per Region, so batch file-arrival events rather than firing one job per file at any real volume.
Only once ingestion actually runs does the two-strategy chunking split from above happen — and only then is the lease searchable at all.
The takeaway
If you’re building RAG on top of any real document set, the question isn’t “should I chunk?” — it’s “does this document set actually have one shape, or several?” Kestrel’s mixed archive (clean template + scanned legacy contracts) is a common real-world pattern, not an edge case. Check your own most-asked-about documents (contracts, policies, runbooks) for real structural nesting before picking a single chunking strategy for everything.
And don’t stop at the chunking decision — the upload path and the sync trigger are just as real a part of the design as the search itself. A developer implementing this needs all three: how documents arrive, what wakes up the ingestion pipeline, and how the live query actually answers.
References
- Amazon Bedrock Knowledge Bases — chunking strategies
- Sync your data with your Knowledge Base (ingestion)
- Build and deploy an automatic sync solution for Amazon Bedrock Knowledge Bases — AWS’s own documented pattern for the S3-event/EventBridge sync trigger
- AWS — Knowledge Bases for Amazon Bedrock now supports advanced RAG capabilities
Part of the AWS Scenario Series — real industry problems, reasoned through on a whiteboard before they’re built.

💬 Comments