Half of the requests we receive for "an AI feature" do not need a model. They need a rule someone has not written down, a database query nobody has run, or a lookup table that has lived in a spreadsheet for years. The other half do need a model, and the difference between the two halves decides whether the project takes three weeks or three months, and whether it costs anything to run afterwards.
This article is the test we apply at the start of an engagement. It is deliberately unglamorous. Its purpose is to make sure a model is built only where a model is the cheapest correct answer.
Start with the decision, not the data
Every machine-learning project is a decision that somebody makes today, made faster, more consistently or at a scale a person cannot manage. Write that decision down in one sentence before anything else: which invoices to hold for review, which support tickets go to which team, what a delivery will cost, which listings are duplicates.
Then ask who makes it now and how. If the answer is "Priya looks at three fields and applies the policy", the first candidate is not a model. It is Priya's policy, written as a rule, with the three fields available in the system. Most operational decisions in a mid-sized company are of this kind, and a rule that encodes them is cheaper, explainable and correct on day one.
The four questions
Can a person state the rule? If an expert can explain how they decide in a handful of conditions, write those conditions as code. A rule engine, a decision table or a few well-named functions will do. The result is auditable and changes when the policy changes, with no retraining.
Does the answer already exist in the data? A surprising number of "prediction" requests are lookups. The expected delivery time is in the carrier's history; the likely spend is in last year's orders. A query, a join and a report answer these without a model. Data engineering, not machine learning, is the discipline they need.
Is the pattern too complex or too variable for a rule? This is where a model earns its place. Text that has to be classified by meaning, images, fraud patterns that shift every month, demand that depends on dozens of interacting factors, and recommendations across thousands of items are all cases where a person cannot state the rule and a query cannot find it. Here the labelled history is the asset, and a model is the tool.
Is being wrong sometimes acceptable? Models are wrong some of the time by design, and the rate is measured, not eliminated. If a wrong answer costs a refund, a model with a review queue for low-confidence cases is fine. If a wrong answer is a regulatory breach or a safety issue, the model can propose but a rule or a person must decide. Deciding this up front shapes the whole design.
What a model costs after launch
The build is the smaller part of the cost. A model in production needs data pipelines that keep feeding it, monitoring for drift as the world changes, retraining on a schedule, an evaluation set to prove each new version is better, and someone who owns all of that. A rule needs none of it. When the four questions point to a rule, the saving is not only the build; it is every month afterwards.
When they point to a model, budget for the running costs explicitly: compute for inference, storage for features and predictions, the monitoring stack and the retraining cycle.
A worked example
A logistics company asked for a model to predict which shipments would be delayed. Discovery found that 70 percent of delays came from four causes a dispatcher could name: a carrier with a known backlog, a destination pin code outside the service area, a missing customs document, and a pickup booked after the cut-off time. Those four became rules, in a week, and caught most of the delays with no model at all.
The remaining 30 percent had no stated cause. For those, the history of a year's shipments, labelled with actual delays, trained a classifier that flagged the risky ones for a dispatcher to check. The model handled the part a rule could not, the rules handled the rest, and the running cost was a fraction of a model for everything.
Where language models fit
Large language models change the second question more than the others. Tasks that used to need a custom model, such as extracting fields from documents, classifying free text or drafting a summary, can now be done with a general model and a good prompt, evaluated against your own cases. That lowers the cost of the "model" answer for language tasks and makes an evaluation set the first deliverable, as the companion article on testing an LLM feature sets out. It does not change the first question: if a person can state the rule, the rule still wins.
The checklist
Before anyone proposes a model, the answers to these should be on one page:
- The decision, in one sentence, and who makes it today.
- Whether an expert can state the rule in a handful of conditions.
- Whether the answer already exists in data a query can reach.
- What a wrong answer costs, and whether a review queue is acceptable.
- What labelled history exists, how much, and how clean it is.
- What the running cost would be, and who would own the model.
If the first two answers are yes, the project is a rule and some data work. If they are no and the last two are answered, it is a model with a budget behind it.
How to bring us the problem
Describe the decision, who makes it now, how often, what it costs to get wrong, and what history you have. That is enough for a scoping call, and by the end of it you will know which of the four questions your problem answers. The AI and machine learning page describes what the build involves when a model is the right answer, and the proposal form is the quickest way to start.


