Artificial Intelligence

Machine Learning Consulting: What Happens After the Model

October 1, 2026 16 min read yatin
machine learning consulting

Machine learning consulting covers scoping, data, model building and deployment — but most of what you pay for over two years comes after the model ships: monitoring, drift detection, retraining, validation and governance. A good engagement prices and staffs that phase up front, and says in writing who owns the model and its data.

Models degrade by default. When researchers tested four standard model types on 32 datasets from healthcare, transport, finance and weather, they observed temporal degradation in 91% of cases. Yet the Reserve Bank of India’s 2025 survey of the entities it supervises found that while 37% of respondents retrain their models periodically, only 21% monitor for data or model drift, and just 14% monitor performance in real time.

That gap — between how models behave and how they are looked after — is what this guide is about: what the lifecycle involves, what it costs, what regulators now expect, how the commercial terms should work, and what to ask before you sign.

The part everyone quotes for

Every machine learning proposal covers the build: scoping the problem, preparing the data, engineering features, training candidate models, evaluating them against a baseline and deploying the winner. It is real work and it deserves a proper budget. It is also the part vendors are most comfortable pricing, because it ends.

It also takes longer than most plans assume. In Gartner’s survey of 644 organisations, only 48% of AI projects made it into production, and the average project took eight months to get from prototype to production. An older survey by Algorithmia found that 64% of organisations took a month or longer just to deploy a model once it was built.

The reason is structural, and it was described a decade ago. Google engineers’ paper on hidden technical debt in machine learning observed that only a small fraction of a real-world machine learning system is the machine learning code; the rest is data collection, verification, feature extraction, serving infrastructure, configuration and monitoring. The same paper warned that it is common to incur massive ongoing maintenance costs in real-world systems. A proposal that is mostly about the model is mostly about the small box.

The part nobody quotes for

After deployment, a model needs six kinds of ongoing work. None of them is optional for a model that makes decisions that matter, and most proposals mention them in a sentence at the end, if at all.

The ongoing work, and what drives its cost

Activity How often Who does it What drives the cost
Performance monitoring Continuously, with daily or weekly review An ML engineer Almost entirely people — the compute is trivial
Drift detection Daily to weekly checks on inputs and outputs An ML engineer Time spent triaging alerts, most of which turn out to be harmless
Retraining On a schedule or when a trigger fires A data scientist and an ML engineer New labelled data, compute, and re-running the evaluation
Re-validation and approval After every retrain, and at least annually for regulated models An independent validator Documentation, testing and sign-off time
Incident response When something goes wrong Whoever is on call How quickly the problem is found and how easily the model can be rolled back
Documentation and inventory At every change The model owner Discipline more than hours — and it is what auditors ask for first

The tooling is cheap. AWS’s own SageMaker pricing example prices a daily monitoring job at about $2.31 a month in compute. The people are not. Glassdoor India puts the average machine learning engineer at ₹9 lakh a year, a senior machine learning engineer at ₹18 lakh and a senior data scientist at ₹20 lakh.

As a planning figure, budget a quarter to half of a senior engineer’s time for each family of production models — roughly ₹4.5–10 lakh a year at those salaries — more for models that make regulated decisions, less for internal ones. The exact share depends on how often the data changes and how much is at stake when the model is wrong. Put it in the business case before the build is approved, not after the first incident.

How a model degrades, and how you find out before your customers do

Models rarely fail suddenly. They get slowly worse, and the first person to notice is often a customer or a branch manager. There are five usual causes, and each has a signal you can monitor.

Why models degrade, and what to watch

Cause Example Signal to monitor
Data drift — the inputs change A lender enters a new region, and applicant profiles shift Distribution of each input compared with training data
Concept drift — the relationship changes The same income and bureau score now carry a different risk after a policy or economic change Accuracy on recent labelled outcomes, as they arrive
Upstream change — the pipeline changes A source system renames a field or changes a unit, and a feature silently becomes zero Null rates, ranges and schema checks on every feature
Seasonality Festive-season demand, harvest cycles, financial year-end Performance compared with the same period last year, not just last week
Feedback loops A model’s own decisions shape the data it is retrained on — declined applicants never produce repayment data Performance on a small, controlled holdout group

Concept drift is the hard one, because you only see it when outcomes arrive — a loan’s repayment behaviour takes months to show. That is why good monitoring watches the inputs and the model’s own outputs daily, and uses them as early warnings while waiting for the ground truth.

The tools are mature, and several are open source: Evidently AI, NannyML (now part of Soda), whylogs, Arize with its open-source Phoenix, and Fiddler, alongside the monitoring built into SageMaker and Vertex AI. Choosing one takes an afternoon. Deciding who reads the alerts, and what they do next, is the actual work.

Retraining: cadence, trigger and cost

There are two ways to decide when to retrain: on a schedule, or when a trigger fires. Schedules are simple and predictable; triggers are efficient but need monitoring you trust. Most production systems use both — a scheduled review, with triggers that can bring it forward.

Typical retraining patterns by model type

Model type Typical trigger Typical frequency Effort per cycle
Credit risk scoring at NBFCs and digital lenders Performance decline, policy change, new products or regions Periodic review at least annually, plus triggered reviews High — validation, documentation and approval
Fraud detection Drift alerts and new fraud patterns Weekly to monthly Medium
Demand forecasting Seasons, promotions and new products Monthly or quarterly Medium
Document classification and extraction New document types or layout changes When triggered Low to medium
Churn and propensity models Campaign or pricing changes Quarterly Low
Language-model assistants over your documents New documents, and any change to the underlying model Index continuously; re-run the evaluation set on every model change Medium

Retraining is not free and it is not risk-free. Each cycle needs fresh labelled data, compute, a full re-run of the evaluation, and — for regulated models — independent validation and approval before the new version goes live. A retrained model that performs worse than the old one on a segment nobody checked is a common and avoidable incident. Keep the previous version deployable until the new one has proved itself.

Language models change the lifecycle, not the need for it

More machine learning work now involves large language models, and the lifecycle looks different. When the model is a hosted service, the provider can change or retire it on its own schedule. What you own is everything around it: the prompts, the retrieval index, the guardrails and — most important — the evaluation set that tells you whether the system still answers correctly.

That evaluation set is the equivalent of a monitored holdout for a classic model. Build it from real questions with agreed answers, weighted towards the cases that matter most, and re-run it on every change: a new model version, a revised prompt, a re-indexed document store. Monitor what users actually experience between changes — how often retrieval finds the right passage, how often the system declines to answer, how often people correct or ignore it, and what each answer costs. Keep prompts versioned like code, so a regression can be traced and rolled back.

Regulators draw no distinction. The Reserve Bank’s draft model risk guidance applies to all models used by regulated entities, including third-party models and those employing AI and machine learning — which covers a hosted language model inside a customer-facing process. If you need control over when the model changes, running an open-weight model on your own servers is one way to get it.

Governance: what your risk and audit teams will ask for

Indian regulators have moved from principles to specifics. In August 2025 the Reserve Bank published its Framework for Responsible and Ethical Enablement of AI, with 26 recommendations. They include a board-approved AI policy for regulated entities and governance across the whole AI lifecycle. The survey behind it was blunt: of the 127 entities that reported using AI, only 15% used interpretability tools such as SHAP or LIME, and only 18% kept audit logs.

In June 2026 the Reserve Bank went further, releasing a draft of broad model risk management guidance that applies to all models used by regulated entities, including third-party and AI and machine learning models. It was still a draft when this was written — the comment period closed in July 2026 — but its direction is clear, and it is worth building to now. The draft expects:

That last point changes the vendor conversation. If you are an NBFC, a digital lender or an insurer, the consultant who builds your model is a third party in the regulator’s terms, and you remain accountable for what it does. Their documentation, validation evidence and monitoring become your evidence. Contract for them accordingly. And if the model is trained on personal data, the DPDP Act adds obligations of its own around security, breach reporting and erasure.

The commercials

There are four common ways to buy machine learning work. Each is right for some situations, and each has a clause to watch.

Engagement models compared

Engagement model Suits How ownership is usually handled Watch for
Fixed-fee build A well-scoped model with known, accessible data Negotiable — insist that you own the trained model, code, features and documentation Change-request clauses, and a price that stops at deployment
Time and materials Exploratory work where the answer is not yet known Work product usually belongs to the client Open-ended spend without milestones or a stopping rule
Build-operate-transfer You want it run for a period, then taken in-house Transfers at a defined milestone The quality of the handover: runbooks, documentation, training
Managed service You do not want to build a machine learning team The vendor may keep its platform; you should still own your model, features and data Lock-in and exit terms — agree them before you need them

Whatever the model, the running phase should be priced as a line item, with a named level of effort, not left as “support as required”. A vendor who will only price the build is telling you they do not intend to be there for the part that decides whether the model keeps working.

Ten questions to ask before you sign

  1. Who owns the trained model, the code and the features? In writing, including anything built on your data.
  2. What happens to our data? Where it is stored, who can access it, and how it is deleted when the engagement ends.
  3. Is retraining included, and what triggers it? A schedule, a trigger, or both — and who pays.
  4. What is the performance commitment? The metric, how it is measured, on which data, and what happens if it is missed.
  5. What is monitored, and who sees the alerts? Named signals, named thresholds, a named person.
  6. Who validates the model independently? Especially for credit, pricing and claims decisions.
  7. What documentation will we receive? Enough for our own model inventory and our auditors, not just a slide deck.
  8. How do we roll back a bad model? How long it takes, and whether the previous version stays deployable.
  9. What third-party components are inside? Pre-trained models, libraries and services, with their licences.
  10. What is the exit plan? Handover, knowledge transfer and access to everything you will need to run it without them.

A vendor who answers all ten clearly and in writing has usually run models in production for years. One who answers only the first three has usually built models and moved on — and you will discover which kind you hired at the first incident.

How we work

We build machine learning systems to be run, not just delivered. Every engagement starts with a 4-week AI proof of concept on your own data with a success measure agreed in week one, so the business case includes real accuracy numbers rather than hopes. When the build goes ahead, monitoring, drift detection, the retraining process and the documentation your auditors will want are part of the scope from day one, and the data underneath is handled the way we describe in our guide to data engineering consulting.

We stay on to run the models if you want us to, through our secure deployment and MLOps practice, or hand them over with the runbooks, inventory and training your team needs. Built fully custom, or accelerated by our own LLM and inference stack where it speeds delivery — the choice is yours. Our AI consulting work starts earlier, deciding which models are worth building at all, and the models run as easily on your own servers as in the cloud. For the wider programme, see our enterprise AI solutions.

Frequently asked questions

What does a machine learning consultant do?

A machine learning consultant scopes the problem, assesses and prepares the data, builds and evaluates models, deploys the best one into your systems and — in a complete engagement — sets up the monitoring, retraining, validation and documentation needed to keep it working in production.

How much does machine learning consulting cost?

The build is usually the smaller part over two years. Budget separately for running each production model family — roughly a quarter to half of a senior engineer’s time, about ₹4.5–10 lakh a year at current Indian salaries — plus retraining, validation and any regulatory documentation.

How long does it take to build and deploy a machine learning model?

Weeks for a proof of concept on accessible data; months for production. Gartner’s survey found the average AI project took eight months to go from prototype to production, and only 48% of projects got there.

How often should machine learning models be retrained?

It depends on how fast the data changes. Fraud models may need weekly or monthly retraining; forecasting models monthly or quarterly; credit models at least an annual review under a board-approved policy. Combine a schedule with drift-based triggers, and validate every retrained version before it goes live.

What is model drift?

Model drift is the gradual loss of a model’s accuracy because the world it predicts has changed. Data drift is when the inputs change; concept drift is when the relationship between inputs and outcomes changes. In one study of 128 model-and-dataset pairs, 91% showed degradation over time.

Who owns a machine learning model built by a consultant?

Whoever the contract says. Insist in writing that you own the trained model, code, features and documentation, and that your data is deleted or returned at the end. For regulated entities, remember the regulator holds you accountable for the model’s outcomes whoever built it.

What does the RBI expect for AI and machine learning models?

Its June 2026 draft guidance expects a board-approved model risk framework covering AI and machine learning models, independent validation including of third-party models, ongoing monitoring for data and concept drift, annual risk-tier reviews and a full model inventory. Check the current status before relying on it.

Y

yatin

Enterprise AI team at AIVeda.

← Previous

AI Strategy Consulting: How to Build a Roadmap That Pays