The Price of a Question Why this is buildable now

The interesting number in AI right now is not on a benchmark.

For most of the past three years, the answer to "could a community health worker simply ask a machine?" was yes, and the arithmetic doesn't work. The model was good enough for the question in front of her: a drug interaction, a referral threshold, what the guidance says about a child that age. And the cost of asking was more than any budget within a hundred kilometres would absorb, for every question, forever.

That stopped being true quietly enough that most people haven't updated. Below a certain cost per question you stop metering: you stop building the billing, you stop rationing by price, and a health worker with a question at four in the afternoon asks it instead of deciding it isn't worth it. Almost everything Grow is depends on being on the correct side of that line.

This is what put us there, what we think the technology is for once you are, how a question actually gets sent to a model, and what none of it has measured yet.


Contents

Sections I to III are the argument, IV to VI the mechanism, VII the list of what none of it has earned the right to claim. V is where the strongest published argument against us gets its say.

I. The floor fell out of the price Open weights, a small end that caught up, and a ladder spanning an order of magnitude.

II. What this technology is for out here The expert is already there. What's missing is somebody to ask before the moment closes.

III. The bill is denominated in reach A fixed pot means a wasted token is a question somebody else doesn't get to ask.

IV. The ladder, the rule, and the shadow Three rungs, five settings, two structural tests, and a mode that measures the cheap one.

V. Why there is no clever router Four papers, including the ones that beat us, and the production router that scored below every model it chose between.

VI. Cache before you route A cache hit costs nothing and loses no quality. A cheaper model can't say that.

VII. What we haven't proved What is live, what is unconfigured, and what has no traffic behind it at all.


Section I

The floor fell out of the price

Run this programme in 2023 and you get one flagship model, one vendor, one rate card, and a bill that scales with how useful the thing is: a funding model where success is what kills you. Every district added is a permanent rise in the monthly invoice, and every health worker who trusts it enough to ask a second question makes it worse. You end up metering, and then deciding whose questions are worth answering.

Three things changed between then and now, and they compound.

Open weights broke the monopoly on the bottom of the ladder

Since 2024 the competent small models have not been one company's product to price. Qwen, Llama, Gemma and Mistral all publish weights under licences that let somebody else run them, and that does two things at once: it drags the hosted rate cards down, because a vendor pricing the cheap end is competing with something anyone can host, and it means the model can eventually be run by somebody other than the vendor: in-country, on a partner's own hardware, and one day on the handset in the field1.

The small end got good

Look at what a frontline worker actually types. A drug name and a dose. A pest on a leaf. What the guidance says about a child under five. Explain this again in the language the class thinks in. A great deal of it is lookup, translation, restatement and arithmetic: work that stopped needing the frontier model some time ago. What it needed was for somebody to check that rather than assume it, which is what Sections IV to VI are about.

And one vendor's ladder now spans an order of magnitude

Grow's service is configured against Alibaba Cloud Model Studio's Qwen ladder: a flagship (qwen3.7-max), a middle rung (plus) and a cost-optimised one (flash), all behind a single key, each with its own published price per million tokens. Every major vendor publishes the same shape of card (Alibaba, Google, OpenAI, Anthropic, Mistral), so the ratio between the top and bottom of a ladder is a thing you can check for yourself in ten minutes, at today's prices rather than the ones we happened to see. On the cards we have priced against, the bottom rung sits at roughly a tenth of the top2.

A tenth is not an optimisation. A tenth is the difference between carrying the cost for a district and carrying it for a country.

Relative price per token flagship every turn, today mid unused by the rule cost-optimised short first turns schematic, not a rate card
Figure 1. The ladder as the service sees it: three named rungs, one switch. Bar widths show the shape, not a quote: the real card is configuration, set per model, and a model with no price recorded produces no cost figure rather than a zero.3
Section II

What this technology is for out here

Most of the capital in this industry is chasing the replacement of expertise. It is a coherent bet and it is not ours. The framing that comes with it (the model knows better than you, so the model must be enormous, and success looks like fewer people in the loop) inverts almost every fact about the setting we work in.

A farmer with thirty years on the same land is not short of judgement. A nurse who knows every household in the catchment is not short of judgement. A teacher running four year groups in one room is not short of judgement, and would tell you so. What they are short of is a specific fact, in their language, before the moment closes: what is on this leaf, what rate for a plot this size, what the referral guideline says, how to put this so a nine-year-old follows it.

The knowledge nearly always exists somewhere. It just tends to arrive after the planting window shut, after the class moved on, after the child got worse.

That is a much easier problem than the one the industry is funding, and a model that would have been frontier-class two years ago solves most of it. It is also a problem where the value is in arrival rather than brilliance: an answer in ten minutes that is 90% as good as a specialist's beats a specialist's answer at the next scheduled visit, because by then the decision has been taken without either. That is why the cost of a question is not an operational detail here: a tool that is free to ask gets asked the second question, and the second question is where most of the value was.

The three commitments this rests on

It does not decide about people. On health it explains, translates and recalls what the guidelines say. It does not diagnose, score eligibility, or order a queue. The health worker does, as they already do.

It is not trained on what anyone asks it. Questions and answers are not training data. Not ours, not anyone's, not aggregated, not "for research".

The judgement stays where it already is. Every deployment starts with an organisation that has been doing the work for years. They set the brief and keep the judgement; the handover is the deliverable, not a milestone at the end of a plan.

Section III

The bill is denominated in reach

Here is the difference that reorganises the engineering.

A commercial AI product spends tokens against margin. Spend more per query and you make less per customer; the constraint is a gross margin somebody will forgive you for missing this quarter. A programme that has bought a block of capacity up front is somewhere else entirely. The pot is fixed. Every token spent on a question that did not need the flagship is not a smaller margin; it is a question some other health worker does not get to ask.

Tokens are not cost of goods here. They are reach.

Once you see the bill in that currency, decisions that look like penny-pinching turn into fairness questions, and some of them reverse. Sending every turn to the flagship stops being the safe, generous default and becomes a choice to serve fewer people very slightly better. Sending every turn to the cheap rung to stretch the budget stops being thrift and becomes a decision to quietly give somebody a worse answer to a question you cannot see, in a language you may not read, with no way for them to check it.

Neither is obviously right, and neither should be settled by whoever is feeling more confident that week. It should be settled by measurement, which is what the rest of this is about.

Section IV

The ladder, the rule, and the shadow

Behind one endpoint there is a ladder of models and a rule that picks a rung. The app never learns which one answered: the service reframes whatever the provider streams into a small vocabulary of its own, so swapping a rung, or the entire provider, is a configuration change rather than a release every handset in the field has to install. On a shared phone in a place with intermittent connectivity, being able to change the model without asking anyone to update is the difference between reacting and not.

ModeWhat happens
off Every turn to the flagship. The default in the code, and the setting that cannot give anyone a worse answer.
static The structural rule below: short first turns go to the cheap rung, everything else goes up. This is what production is running.
shadow Serve from the flagship as normal; on a sampled fraction of turns, also ask the cheap rung and log both.
cheap / mid Pin everything to one rung. For a load test, or a cost emergency where the alternative is being switched off entirely.

The rule, in full

Two tests, both structural, both erring upward. A follow-up goes to the flagship: it is where a cheap model has the most to lose, because it has to hold the whole thread, and the failure mode is quietly dropping what was established three turns ago, which is invisible to somebody who cannot check the answer against a specialist. A request longer than 400 characters goes to the flagship too: length is a proxy for somebody pasting something complicated. Everything left over is a short first question, and that is what the cheap rung takes.

That is the whole rule, and those are the real numbers. As this is published the service reports mode static, a flagship of qwen3.7-max, a cheap rung of qwen3.6-flash, a ceiling of 400 characters and a shadow sample rate of zero, which you can read off its own health endpoint rather than taking from us.

a turn follow-up? turn 2 or later over 400 chars? the whole request yes yes flagship no cheap rung two tests both structural, both erring upward
Figure 2. The whole rule. It reads two numbers the service already records (how many turns deep the conversation is, and how long the request was) and nothing else. No model runs to make this decision.

Shadow mode, or: measure before you switch

None of this is worth anything if the cheap rung is worse in a way that matters, and there is one honest way to find out: ask it the same questions and compare. Shadow mode answers every user from the flagship, as normal, and on a sampled fraction of turns quietly asks the cheap rung the same thing. Both rows land in the ledger sharing a turn id, so the pair joins. The shadow answer is never shown to anyone and never stored4; what it leaves behind is what it cost, how long it took, and something to compare.

It is sampled rather than universal because on a fixed pot a shadow call is a real answer somebody else does not get. At one turn in twenty the comparison set grows steadily and the bill barely moves; at every turn you pay twice for every question to learn what a fiftieth of them would have told you. The rate in production is currently zero, which is the uncomfortable half of Section VII.

Section V

Why there is no clever router

The obvious build, the one suggested in every conversation about this, is a model that reads each question and decides where to send it. Start with the strongest evidence against us, then the case.

Routing works. Ding et al. route with a DeBERTa model over the query alone and make up to 40% fewer calls to the large model with no measured quality drop5. Zhang et al.'s MTRouter, choosing a model per turn, beats GPT-5 on ScienceWorld, 53.8 against 48.4, while cutting total cost by 58.7%7. Neither of those is a straw man, and the second is more ambitious than anything Grow does.

So the question is not whether routers work. It is what happens to one when it meets a district it has never seen, in a language it was not trained on, with no way for anyone there to tell it went wrong.

What happened when a production router met an unfamiliar domain

The same MTRouter paper benchmarks a commercial routing service, OpenRouter's automatic API, on the same tasks. It scored −26.4 on ScienceWorld, against 48.4 for GPT-5 alone and 53.8 for their own router. Not merely worse than the best model: worse than every one of the six models it was choosing between, including the cheapest. The authors' explanation is the part that matters here: it underestimated task difficulty in an unfamiliar domain and over-relied on the lightweight models.

That is the failure this programme has to design against, and it is not a hypothetical we invented. A router mis-reads the difficulty of questions it wasn't trained on, sends them to the cheap model, and returns confident answers nobody flags. Now put that in a clinic where the question is in Bengali and there is no specialist within a day's travel to notice.

1. The savings have to survive the router itself

The JAIR survey of routing strategies states the paradox plainly: routing infrastructure designed to reduce costs "may ultimately consume more resources than it saves in the long term", which it calls a clear case of technical debt, and concludes that operators must evaluate whether the operational savings justify the cost of building, training and maintaining the router6.

Being precise, because this argument is conditional and we have overstated it before: the cheapest routers, similarity and clustering methods, are genuinely low-compute, and the survey credits them with generalising to new candidates. The expensive ones are the trained ones. So the tax is real for exactly the class of router that would be any use to us, and the class that would be cheap is the class that degrades in the regime we operate in. Ding et al. measured that regime: held to about 1% quality loss, their large-gap pair yields a cost advantage on test of 4.4% to 5.1%, and their untransformed routers there do only marginally better than routing at random.

2. The annotation bottleneck, which scales the wrong way

Training a router means labelling: for each query, what each candidate would have answered and how good it was. The survey works the arithmetic through on a public benchmark: 112 candidate models over roughly 16,000 questions is about 1.8 million generations, which then have to be evaluated. And the cost compounds in the direction that matters to us: adding one candidate means generating and evaluating responses across the entire query set. Even similarity-based routing cannot escape it.

MTRouter's authors report what that cost them, and it is the only figure in this literature denominated in the currency this essay uses: 1,291 training instances, 29,693 logged trajectories, 515,221 turns, about $1,620 in one-time collection. That is what one well-resourced team spent to build the thing we are declining to build, before any of it is maintained, and on a benchmark rather than on a district's real questions.

Grow's entire deployment story is that a rung, or the whole provider, is a configuration change, not a release. A component that has to be re-annotated every time we change a model is the wrong shape for that, and $1,620 of inference spent on labelling is $1,620 not spent on answering anybody.

3. A router is a control plane, and a control plane is a target

This one we had not thought of, and it is the survey's: because the router decides where everything goes, it can be attacked directly. Adversaries can craft inputs that consistently trigger selection of the most expensive model, inflating operational cost without necessarily affecting the answers anyone sees. The survey calls this an LLM control plane integrity attack, alongside poisoning public routing benchmarks to plant triggers.

For a commercial service that is a margin problem discovered on an invoice. Here, spend is reach: an attacker who can inflate cost takes questions away from health workers, quietly, without ever producing a wrong answer for anyone to notice.

And our rule does not escape this. It is more trivially gameable than a learned router, not less: anyone can pad a question past 400 characters or send a second turn, and every request lands on the flagship. An earlier draft of this page called that a safe failure because it fails towards the better answer. That was wrong, and it was wrong in the exact way this section has just finished describing: forcing the expensive model is the control plane attack, and the fact that each individual answer stays good is what makes it hard to notice.

What contains it is not the routing rule. It is the per-account daily cap in Section IV, which bounds what any one caller can spend in a day whatever the router decides. That is the honest division of labour: routing decides what a question costs, and the cap decides what an attacker can take. A structural rule is simply easier to reason about when you are working out which of the two is protecting you.

4. Training one means training on what people typed

This is the argument no measurement can touch. A learned router is learned from traffic. Ding et al. train on a corpus of real user instructions, MTRouter on nearly thirty thousand logged agent trajectories. Neither paper raises consent, and for their purposes neither has to. This programme has promised in writing that questions are not training data. That promise is worth more to the people using it than a few percent of a budget is worth to us, and it is not a decision an engineer makes quietly on a Thursday.

And the language argument, stated properly

The version we used to make, that an English keyword list fails in Bengali, is a straw man against a learned router, which reads meaning rather than keywords. The real objection is distributional, and now it has three pieces of published support. Ding et al. measure transfer decaying towards the random baseline as the correlation between model pairs weakens. MTRouter's ablation shows that stripping routing history drops ScienceWorld from 53.8 to 44.6, so the query-only configuration, the one we would be able to build, is the weak one. And OpenRouter's −26.4 is what a production router does in an unfamiliar domain.

The app ships twelve languages today: English, Spanish, French, Portuguese, Arabic, Hindi, Bengali, Urdu, Swahili, Indonesian, Tagalog and Chinese. A router learns the traffic that exists, which at the start is disproportionately English, and the other eleven are the unfamiliar domain. The decay would be invisible, because a question wrongly called easy returns an answer rather than an error.

A rule that fails invisibly on the people with the least recourse is worse than no rule at all.

What the literature says for the boring rule

Three things, none of them ours.

Conservatism is the finding, not the naive position. MTRouter beats the router it is compared against by switching models less: after an error it stays with the same model about 90% of the time on ScienceWorld, against 38% for Router-R1. Its authors are explicit that multi-turn routing is not "switch more".

Switching has a cache cost. The same paper notes that frequent switching lowers prompt-cache hit rates in multi-turn settings and raises the effective cost of serving long histories. Our rule sends every follow-up to the flagship and therefore never switches mid-conversation. We adopted that for a quality reason; it turns out to have a cost reason too.

Per-query routing controls a hard budget only indirectly. LinkedIn's batch-level work makes the point that per-query routing trades cost against quality through a weighting term, which makes strict budget constraints difficult to enforce. Their answer is to optimise over a whole batch under an explicit cap8. A fixed pot is exactly a hard constraint. Our cap is enforced separately from routing, per account and per day, for that reason.

What we no longer cite, and why

Earlier versions of this page leaned on two unpublished numbers of our own: a query-only router predicting the cheap model's failures at ROC-AUC ≈ 0.5, and a static policy beating a trained classifier 0.89 to 0.76. The second was never a fair comparison: the two policies moved different fractions of traffic, and the correct way to compare is at a matched cost advantage, as Ding et al. do.

Both are now redundant, because the published record already says it and says it checkably. The survey records the original RouteLLM strategy beating random routing by 56% to 51% in one experiment, a separate RouteLLM implementation not outperforming random routing at all, verbalised-confidence routing performing no better than random, and AutoMix failing to beat random routing on RouterBench. It also records what happens off-distribution in the supervised case: domain routing that reaches 77–97% accuracy in specialised areas falls to as low as 53% on "other" queries. Every one of those is a citation a reader can follow. Ours were not, so they are withdrawn rather than defended.

Section VI

Cache before you route

One thing goes ahead of all of this, and it is why the ladder may end up switched off again.

Any given service has twenty questions that eat most of its week. The same drug, the same pest, the same threshold, asked by different people in the same district within days of each other. A cache hit costs nothing and loses no quality, precisely what a cheaper model cannot promise. So the first script to run against real traffic counts duplicates, not tokens. If the repeat rate is high, caching dominates routing outright and the correct decision is to leave the ladder alone.

Section VII

What we haven't proved

The programme's standing promise is that you will be told which parts are running, which are prototypes and which are still intentions, including the awkward ones. Applied to this essay, on the day it was published:

The rule is live, and nothing has measured it. An earlier draft of this page said routing shipped switched off. That was wrong: production is running static, so short first turns are already going to the cheap rung. What has not happened is the measurement: the comparison that would say whether those answers are as good.

Shadow mode is not sampling. The rate is zero, so the apparatus described in Section IV is built and idle. Nobody is currently collecting the pairs that would settle it.

No prices are configured, so nothing is costed. The ladder exists to produce a cost per turn and the service is recording none, because the rate card was never set in the deployment. The column is null rather than wrong3, which is the right failure, but it means the saving this page is about is, at this moment, unquantified by us as well as by you.

There is barely any traffic to measure anyway.9 Every projection in this space is a story until it meets a real week of questions from real people, and ours hasn't yet. Some two thousand words of position sit on top of that; treat the ratio as the honest state of things rather than as confidence.

The models are hosted, so asking needs a connection. On-device is where this is going and it is the single biggest thing that would change how far it reaches. It is a goal. This page will not imply it already works.

Cheaper is not the objective. The objective is the largest number of good answers per unit of capacity. If the measurement says the cheap rung is worse in a way that matters, the ladder goes back off and the flagship is simply what this costs, at which point the honest conversation is about funding, not engineering.

What you can check

The rule is stated in full in Section IV (two tests, one threshold, no model, nothing withheld), and the two arguments it rests on are in Section V and ask you to take nothing on trust. The service's own health endpoint reports the mode, the rungs and the threshold, so the claims about what is running are checkable without us.

What you cannot do is read the code. The service has no dependencies at all, because a thing that holds a credential and spends a budget ought to be small enough for one person to audit in an afternoon, but that is a description of how it is built, not an invitation. The source is not public today.

Notes

  1. Not how it works today. Every question in the Grow app leaves the phone, goes to Grow's own server, and on to a model hosted by another company; the privacy page says exactly where. Open weights are why the direction is credible, not what is running.
  2. "Roughly a tenth" describes the shape of the published cards at the time of writing, not a rate anyone has quoted us. The model names above are checkable and are what the service is configured against; the ratio is deliberately not printed as a figure, because a number on this page would be stale before a partner read it.
  3. By design each turn is priced from the card for the model that actually served it, at the moment it served it, and a model with no price configured yields no figure rather than a zero: zero is a number somebody will later sum, and a column of zeros reads as a free service instead of an unconfigured one. In this deployment no card is configured at all, so that safe failure is the one currently happening.
  4. No message row, no conversation counter, and it is excluded from the per-account daily cap: that spend is the operator's experiment, not the user's usage. The cap itself fails open: if the ledger cannot be read the question goes through, because a health worker mid-question is the wrong person to absorb a database outage.
  5. Dujian Ding, Ankur Mallick, Chi Wang, Robert Sim, Subhabrata Mukherjee, Victor Ruhle, Laks V. S. Lakshmanan and Ahmed Awadallah, Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing, ICLR 2024. Router is DeBERTa-v3-large (300M) over the query alone; quality is BART score; "cost advantage" is the fraction of queries routed to the small model rather than a sum of money. Their code is at github.com/m365-core/hybrid_llm_routing. The figures quoted here are from its Table 1 (cost advantage against quality drop by model pair), Table 2 (router latency), Table 3 (thresholds calibrated to ≤1% drop) and Appendix A.4 (generalisation to unseen model pairs).
  6. Clovis Varangot-Reille, Christophe Bouvard, Mathieu Ciancone, Antoine Gourru, Marion Schaeffer and François Jacquenet, Doing More with Less: A Survey on Routing Strategies for Resource Optimisation in Large Language Model-Based Systems, Journal of Artificial Intelligence Research 86, Article 14 (2026). The technical debt paradox is §5.2.1, the annotation bottleneck §5.2.3, the control plane attack §5.2.5. The near-chance results are in §4.1.3 (RouteLLM at 56% against 51%), §4.3.2 (a RouteLLM implementation not beating random), §4.4.3 (verbalised confidence) and §4.4.5 (AutoMix on RouterBench); the 77–97% to 53% domain-routing figure is §4.2.2, after Simonds et al. 2024.
  7. Yiqun Zhang, Hao Li, Zihan Wang, Shi Feng, Xiaocui Yang, Daling Wang, Bo Zhang, Lei Bai and Shuyue Hu, MTRouter: Cost-Aware Multi-Turn LLM Routing with History–Model Joint Embeddings, ACL 2026. The OpenRouter score, MTRouter's own results and the episode-level baselines are Table 3; the routing-history ablation is Table 4; the switching and recovery rates are Figure 4; the collection cost is §3.3.
  8. Jelena Markovic-Voronov, Kayhan Behdin, Yuanda Xu, Zhengze Zhou, Zhipeng Wang and Rahul Mazumder, Robust Batch-Level Query Routing for Large Language Models under Cost and Capacity Constraints, CAIS 2026.
  9. At the time of writing the usage ledger holds tens of rows from a handful of days, a pilot's worth of traffic, not a service's.

If you think this is wrong about the prices, about what the work actually needs, or about who should be carrying the bill, then hello@fantail.ai reaches a person, and a specific disagreement gets a specific answer.