index/engineering/time-to-answer/One Call Instead of Two
engineeringtime-to-answer

One Call Instead of Two

The discount case from part 3, answered again through a single tool call instead of two manual steps: the joining moved into code, the reading moved into a document, and asking again has no per-token bill on it.

tags: mcp, debugging, local-models, feedback-latency

Part 4 of 4.

This piece is the story. The tooling behind it gets its own post: the ETL, the MCP tool, and the knowledge file, including how I write and test them. The payloads and knowledge entries quoted here are simplified samples, not the real files.

A customer bought a book one day after the sale ended, and the discount didn’t apply. Same case as the last piece, same dates, same conclusion. This time it goes through the tool instead of my hands: one call in place of two. I’m reusing it on purpose: it’s already on the page in full, and what I want to look at this time is where the thinking happened.

The short version: the joining moved into code, the reading moved into a document, and a small local model applies both.

Part 3 was me doing that by hand: write the comment, let the autocomplete draft the query, run it, highlight the slug, fire the curl call off a keyboard shortcut, read the answer in the Run panel. Two questions, two tools, and one human holding the join between them.

I didn’t rebuild it for speed. Part 3’s chain was fast, and I’d still call it fast. What it produced was a new query every time, and I read that query before running it, every time. That check costs little on any single call and never stops.

The rest of the motive is smaller. Two questions against two different things, with a value carried between them by hand, stays a hassle on the hundredth discount complaint.

What I’d defend is the other part: it’s fun. Production issues are the part of the job, and solving one by typing a sentence instead of writing a query is fun. Fun is what got the tooling built. The ETL and the MCP tool had to be hand-written before any of this worked, and that’s the part that costs effort. For the kinds of issue they cover, every case after that costs me a sentence. A new kind of issue isn’t free: it needs its own ETL and its own knowledge entries, so the effort isn’t paid once, exactly.

The split matters for everything below. The joining is written down once, in code: a small ETL that pulls the discount rows and the checkout answer together and trims them to what matters. The same question fetches the same data the same way. The reading is different. A model reads what came back, guided by a knowledge file, the document I wrote that says how to interpret the shape of the response, and reading isn’t repeatable. Ask twice and you get two wordings, sometimes two emphases. A good part of this piece is about where the fixed half ends and the generated half begins. The last step turns the reading into a message a customer can read.

One call, both questions

The invocation is the question, in the shape the tool takes: a sentence out of a ticket, with the book and the order date in it.

fetch discount by book title "testing is fun" order date 2026-06-01, user says they can't get the discount applied
fetch discount by book title "testing is fun" order date 2026-06-01 with previous discount window 5

The signature already does some work before anything comes back. A discount is a fact about a moment rather than a property of a book, and asking about one without a date is asking something undecidable. Once the date is a parameter, the question has an answer, and it has already been narrowed to one book and one instant, which is half of why the response can be as small as it is. The count comes out of the sentence as well. Two windows is the default, and naming a number overrides it.

Behind that line, the tool does the two things part 3 made me do in two places: it reads the discount rows, and it asks the service what a customer would have been shown at checkout. The ETL combines the two into one payload instead of comparing them, because they answer different questions: the service can say whether the discount shows for this order, but the earlier windows can only come from the database. Both happen inside the one call because the ETL assembles them, and the ETL is strict about what it will assemble. If either lookup fails, the code exits early with an error log and the model never receives a payload. That strictness is how I debug by hand, written down.

The context comes from as many services as the question touches, assembled by the ETL, and the model can still call a detail tool for what the context doesn’t carry, a book by its id. The MCP is a proxy over the services, and what the model can reach is what I wired into it.

It isn’t always one shot, though. The call brings the data into the model’s context, the next thing I ask it to do is read what it just fetched, and at the end I ask for the short version. Whether that takes one turn or several is mine to set, by how well I already know the issue. If it comes back thin or bent, I ask for a simpler or more direct one.

What comes back from that one call is small, and the next parts are about why. First, though, the objection I’d raise myself.

Why the model is local, and why I don’t hand it the spec

Yes… I know, there’s a better-informed reader available. Some of the code I work on is written from a spec now, and more of it will be. A model with a large context window could read that spec and the code together and check one against the other: no ETL, no tool, no knowledge file, nothing for me to keep true. It doesn’t help me, for two reasons.

The first is that on the legacy code I’ve debugged, that document doesn’t exist. Nobody wrote down what the thing was supposed to do, and the code is the only description anybody has. Calling it a description is generous.

The second holds even with the document in hand. The spec is the plan. Production is what happened, and the two drift. Code moves after sign-off, config changes underneath it, and the rows that decide a discount case get written by paths no document describes. When a customer’s discount fails on one order on one date, what I need is what the system did to that row. The document walks me to the code, the code tells me which rows to read, and the rows are the only record of what happened. Neither the spec nor the code has read this order. The row has, and it has no idea what the spec was supposed to say. So I read the rows before I trust either.

That’s the wall. What the system did lives in instrumentation if there is any, and in the database if there isn’t, so something has to read it. Hand that reading to a model and the data goes wherever the model runs. Real orders belong to real people, and I keep that kind of data where it lives. Yes… I know, a large cloud model could do the reading, and probably better than the one I’m using. It doesn’t help. The rows would have to leave to get there, and if it’s a grey area, I stay out of it.

The wall explains the rest of the design. The data stays where it is, so the model runs where the data is: on My Machine. The tool is narrow, because the model must not go looking through the database itself. And the ETL decides what arrives already read, which is what keeps the context small enough for a small model. The wall is why the model is local. The ETL is why a small one is enough.

The answer arrives already read

In part 3 the API answered with discount_value: null, and the null was the problem. Is that a service that looked and found nothing, a config that says no discount, or a bug? I read it by hand and concluded the sale had ended.

Here the property is simply absent. For the case in this piece, the response is:

{
  "book_title": "testing is fun",
  "order_date": "2026-06-01",
  "previous_discount_windows": [
    { "start_date": "2026-02-01", "end_date": "2026-02-28" },
    { "start_date": "2026-05-01", "end_date": "2026-05-31" }
  ]
}

And for an order placed inside the sale window it would be:

{
  "book_title": "testing is fun",
  "order_date": "2026-05-20",
  "book_discount_value": 10
}

book_discount_value is in the response when the discount applies to that book on that order date, and missing when it doesn’t. No boolean, no status: false, no null to weigh. Absence can’t be a failure here: a failed lookup exits the code early with an error log, so the model never sees a payload at all. A missing property means the lookups ran and the discount wasn’t there. The rule is carried by the name of the property that went missing, and the knowledge file handed to the model alongside the data says what to make of it: how to read the shape it arrives in, what an empty result means, and how to take the request that came in with it. One entry looks roughly like this:

## book_discount_value is missing
No discount applied to this book on this order date. This is not an error.
if `previous_discount_windows` exists then the order date falls outside every
campaign this book has had: compare it against the last window, name that window,
and say the sale had ended before the order.
if neither key is present then the book has never had a discount: say that and stop.

A simplified sketch, with the branch for the ordinary case left out. The real knowledge file, and how I write and test it, get their own post.

It’s one choice seen from two sides, and both are about who’s reading. A REST endpoint keeps its shape stable for clients that break on a missing key, so it sends the key and sets the value to null. A response inside this loop has no such client, so it can be shaped for the one reader it has, and the shape can carry meaning: only what applies, named the way the business names it. A boolean flag would have needed a comment somewhere explaining what false meant. A property that appears and disappears carries its own definition.

This is why a 4-billion-parameter model is viable here: the rule it applies was written down once, and the dates it applies that rule to arrive with the question, so what’s left is a comparison rather than a recall.

The payload is the record, the sentence is the interface

Part 3’s answer arrived as a grid, and the thing I was checking wasn’t in any single column of it. The condition spans a few: which discount, on which book, against which date, and whether the window covers that date. Checking that by hand means reading across the row every time and holding the relationship in my head while I do it.

What comes back here is a payload, and the tool would work without the model at all. I could read the JSON myself. What the model adds is the pointing: it states the relationship in a sentence.

The pointing is its whole contribution, and that settles the shape I want whenever I have to choose one. The payload is the part I can check against the database and the endpoint, and that makes it the record. The sentence is what I can act on without lining up dates by eye, and that makes it the interface. I want both, in that order. If you made me drop one, I’d keep the payload, because every value in it is one I can check. What the payload doesn’t do is say it out loud. The order date and the windows are both in there, and the comparison between them is still mine to hold until the sentence states it. Whether that sentence can be trusted is the next section.

The explanation is generated, not stored

What comes back for this order doesn’t say the buyer was a day late. The payload carries the order date and the window it fell outside of, and the sentence that connects them is produced fresh on every run, by the model reading the rules I wrote against the data it was handed. The rules are written down, case by case. What isn’t written down anywhere is this order’s explanation. You can watch that happen. When the first summary comes back too vague to act on, asking again produces a different one from the same data, which is what a reading looks like.

So the two halves don’t deserve the same confidence, and I try to hold them differently. A missing property is a fact the pipeline decided, and I can go and check it. “There’s no discount for this order” is an argument about a rule and a date, and the dates it argues about are the one I typed into the question and the ones the ETL brought along with it. That includes the day count: “one day after” is arithmetic the model does on the fly, so I check it against the two dates. Arguments about rules get checked against the dates before I repeat them.

What the model does

The model is Qwen3-4B-Instruct-2507, running locally because of the wall above. It’s small because local inference is the constraint. I’ve run this loop on models at 12, 20 and 30 billion parameters as well, and the 4 billion one is the one I kept.

The bigger and denser ones need more compute per answer and draw more power, and I found that with enough tweaking of the prompt and the knowledge file, the 4 billion one gets the job done anyway. So I stopped paying for the extra weight.

A model that can’t hold much forces the question of what this needs: anything I don’t cut from the response is context it has to carry, so the limit does some of the deciding before I’ve written the ETL. It also forces me to learn the model, since how it behaves is not optional at this size. Once the knowledge file and the tool carry the load, the model’s size matters much less. Much less isn’t nothing, as the tool-call rate under “Where this stops” shows.

Yes… I know, a 4-billion-parameter model has no business driving a tool loop, and a model this size is shakiest at exactly this kind of work: picking the right tool and filling in its parameters. That part is still the model’s, and it’s where the tenth call goes wrong. What the loop doesn’t ask of it is everything else: which tables, which join, what the fields are called, and what counts as a previous window were all decided upstream. What’s left is the reading. The rules arrive as written entries, and applying the right one to this order’s dates is where the model earns its place.

The model turns the prompt into MCP parameters and calls the tool. The tool calls the resources, transforms what comes back, and sends the output to the model to read. The model isn’t choosing tables, writing queries, or looking around on its own: every tool it can call is one I wired, and the parameters come out of the sentence I typed. The tool and the knowledge file are mine, so what the model reads has a shape I chose.

The model is switchable. A larger one would probably lift the tool-call rate, and I’d swap it in if the gain justified the cost of running it: more memory for the weights, more time per answer, a longer context to read before the first token comes back.

The knowledge file also lists a remedy for each reason an absence can have; the sketch above leaves them out. None is applied. Nothing in the chain can apply one, by design: a fix fired automatically at a production issue is more risk than I’m willing to hand over, so the loop stops at telling me what it thinks happened.

What the machine bought

Part 2’s argument ran one level below this one. There the toll was per check: in the old test loop a run cost a trip out of the file, so the rational edit was to save up changes and spend one run on all of them. Once the trip was gone, a run took seconds and rerunning for each change was affordable. Here a run takes minutes, so the price sits on the attempt, and attempts are what build the part of the work nobody sees: the ETL and the knowledge file, which decide what the model gets to read. That work is per kind of case. Each kind brings its own joining and its own rules for the entry to read, and the sentence is what’s left once they’re written.

Attempts here aren’t free, and I won’t write them up as though they were. The machine cost money, the power costs money every time it runs, and the license is mine. What isn’t there is a meter on the question itself: no per-token bill, no rate limit to wait out, no queue that isn’t my own machine warming up. A hosted model prices the same attempt by the call and by the token, on a bill that arrives every month. Mine mostly came in one piece, before I asked anything; the electricity is the only part that scales with how often I ask. For a loop whose whole point is asking again, that difference decides how often I ask.

So I change a rule in the knowledge file, run the same question again, and read what comes back. The ETL got strict the same way: I wrote it, ran it, and rewrote it until the response stopped carrying anything I’d have to ignore. The filter came later. Testing happens in two stages, and the order is the point. First I call the tool straight against the MCP server, no model in the loop, which tells me whether my own code works against real data. Only then does the model get the assembled chain, and by then the code has already passed.

Yes… I know, none of that makes an answer fast. It still takes a minute or two, sometimes five. Most of that wait is inference: the bigger the context, the longer it takes before the result shows, which is one more reason the ETL cuts the response down. A reasoning phase adds to it, and even without one, generation is indeterminate by nature, so how long an answer takes isn’t fixed. Debugging by hand isn’t instant either: I’m reading the grid and holding the join in my head, and now the wait happens where I’m not doing that. What the price buys is attempts, and attempts are what the ETL and the knowledge file are made of.

Aimed at someone who didn’t ask

The last thing I ask for is the version a customer can read. Plain, no column names, no property names, no internals: what happened to the order, and what it means for them. A run gives me something like this:

Your order was placed on June 1, one day after the sale on this book had ended, so the discount no longer applied to it.

Then, once I’ve read it, it goes into the email or the support chat, and the reply and the debugging answer come out of the same reading.

I’d underline this part more than any of the mechanics. The old chain ended with me holding an answer, and then a second piece of writing started, aimed at someone who had never wanted the query. Here the reading is already in words, so the message is one request away.

It also changes what being wrong costs, which is the part I watch. A grid that says something odd is a puzzle I get to notice before anyone else does. A friendly sentence reads as an answer, and the person receiving it has no grid to check it against. So the outward version gets to say what the payload supports, and no more than that. I read the customer version against the payload before it goes out.

Where this stops

The ETL is my own debugging judgment, applied before the model sees anything, and that is the ceiling as well as the strength. It omits the things I already know to look for, so it can only omit in ways I taught it to omit, and it can’t surface a class of problem I’ve never looked for.

Real cases don’t always hand you a root cause. Sometimes there’s nothing to reach except symptoms, and then the work is following them across the layers, which is what having more than one tool is for.

The written reading can go stale silently. If the business rule changes and the knowledge file doesn’t, every absence still gets read, confidently, under a rule that stopped being true, and nothing in the chain objects. There’s no failing test for a document that quietly became wrong. Determinism has an edge from the other side too: code that answers the same question the same way every time is also wrong the same way every time. Which is the argument for testing the ETL hard. A fixed thing can be verified once, and the trouble is worth it, because every mistake it makes is a mistake on every question that follows. A query regenerated on each call can be read each time, but never verified once and then trusted. None of this retires part 3’s chain. If the tooling isn’t enough, the question goes back through the manual steps, and there’s no version of this where I’ve given up the ability to check. That includes part 3’s agree-or-disagree comparison between the database and the API: this tool combines the two sources rather than comparing them, so when I need to know whether they agree, that’s still the manual chain. Pick your poison. I’d rather have the one I can verify.

Then there’s the rate. The 4 billion model gets the tool call right about nine times in ten. That’s a feel from using it, not a counted sample. When the tenth goes wrong, it’s usually loud: a failed tool call comes back as an error log, and a retry fixes it. The quiet version is the one I can’t rule out, a reply that answers without its data, which is the failure this whole design exists to prevent. I haven’t counted, so I can’t tell you how often that happens. Either way, the retry is mine: the tool runs for one reader, and when a call comes back wrong I start it again with the prompt corrected.

The clock is the last caveat I’ll keep. I can tell you that an answer takes a minute or two, sometimes five. I can’t tell you how much faster that is, because I never put a clock on the manual chain in part 3. That’s the weakest claim in this piece, and I’d rather leave it weak than dress the number up.

Part 3 ended on the pattern not being the tools but the glue, and on wiring one shortcut so that a value on my screen could answer the next question. I still think that’s right. This is what it looks like when the glue sets: the shortcut becomes an artifact.

Two things moved out of my head. The join between the two questions became code. The reading of an empty result became a document. The model applies both and takes credit for neither. The last thing it does is become a sentence a customer reads.

What I own isn’t a habit my hands can be trusted to remember any more. It’s a file and a piece of code, sitting somewhere they can quietly stop being true. That’s the price of never typing that join again.