LLM

On-Premise LLM Deployment in India: Hardware and Cost

September 24, 2026 19 min read yatin

An on-premise LLM deployment in India starts at roughly ₹13–17 lakh for a single-GPU server that runs a 20–32 billion parameter open-weight model, and rises to ₹3.5–5 crore for an eight-GPU H100 server. It only beats a hosted API above a steady monthly token volume, so that break-even is the first number to work out.

The question has become urgent for Indian enterprises for two reasons. The Digital Personal Data Protection Rules were notified in November 2025, with most obligations phasing in over eighteen months. And the appetite to keep AI workloads close is real: in the Nutanix Enterprise Cloud Index published in March 2026, 57% of the 1,600 executives surveyed said they need to run their infrastructure within a single country, and 82% said their current infrastructure is not fully ready to run AI workloads on-premises.

This guide puts rupee numbers on the decision: what the hardware costs, which open-weight models fit which GPU, what renting GPU capacity inside India costs instead, where break-even lands against API pricing, and what the DPDP Act asks of you once the model runs on your own servers.

What running an LLM on your own hardware involves

An on-premise LLM is four purchases, not one. Budgets that only cover the first one are the reason these projects stall in month three.

For a single-GPU deployment the people cost is roughly the same size as the hardware cost over three years, as the worked example below shows. For an eight-GPU server the hardware dominates. Either way, a deployment with no named owner will be running a year-old model on a year-old driver by its first anniversary.

Sizing: which model fits which GPU

GPU memory decides what you can run. The rule of thumb, from Hugging Face’s own documentation, is that a model with X billion parameters needs roughly 2 × X GB of GPU memory in 16-bit precision. At 8-bit that halves to about 1 GB per billion parameters, and at 4-bit it falls to around 0.5–0.6 GB — Hugging Face’s worked example fits a 15.5 billion parameter model into 9.5 GB. On top of the weights you need headroom for the key-value cache, which grows with context length and with the number of people using the model at once.

Quantising to 8-bit costs almost nothing in quality. Red Hat and Neural Magic ran more than half a million evaluations on quantised Llama 3.1 models and found that every quantisation scheme recovered over 99% of the unquantised model’s average score; the peer-reviewed version of the work describes 8-bit floating point as effectively lossless across all model scales. For most enterprise deployments, 8-bit is the sensible default and 4-bit is a reasonable choice when memory is tight.

Mixture-of-experts models change the arithmetic. Memory is set by the total parameter count, speed by the active count. Llama 4 Scout has 109 billion parameters but only 17 billion active per token, so it needs the memory of a large model and runs at the speed of a mid-sized one.

Open-weight models and the smallest practical GPU setup (8-bit unless stated)

Model Parameters (total / active) Licence Approx. weights at 8-bit Smallest practical setup
OpenAI gpt-oss-20b 21B / 3.6B Apache 2.0 Runs in 16 GB (OpenAI) 1 × L40S 48 GB
Gemma 3 27B 27B dense Gemma Terms of Use ~27 GB 1 × L40S 48 GB
Qwen3-32B 32B dense Apache 2.0 ~32 GB 1 × L40S 48 GB (4-bit for long contexts)
OpenAI gpt-oss-120b 117B / 5.1B Apache 2.0 Single 80 GB GPU (OpenAI) 1 × A100 or H100 80 GB
Llama 4 Scout 109B / 17B Llama 4 Community License ~109 GB 2 × A100 or H100 80 GB
Mistral Small 4 119B / ~6B Apache 2.0 ~119 GB 2 × A100 or H100 80 GB
Qwen3-235B-A22B 235B / 22B Apache 2.0 ~235 GB 4 × H100 80 GB; 8 for long contexts
Llama 4 Maverick 400B / 17B Llama 4 Community License ~400 GB 8 × H100 80 GB

Read the licence before you read the benchmark. Apache 2.0 models can be deployed commercially without conditions. Llama 4 requires a separate licence from Meta only for companies above 700 million monthly active users, which will not affect most Indian enterprises. But some licences are written to exclude exactly the buyers reading this: Mistral Medium 3.5’s modified MIT licence excludes companies whose global monthly revenue exceeds $20 million.

What a GPU server costs in India

GPU prices in India are not published the way cloud prices are. The figures below are retailer listings and distributor estimates as of September 2026, and they move with the rupee and with supply — get three quotes before you budget. When we checked in September 2026, the L40S was listed as out of stock at both Indian retailers we looked at.

Three reference builds and what they cost

Build What it runs well Indicative price in India GPU power draw
A — 1 × L40S 48 GB server 20–32B models for one department or a few hundred users of an internal assistant Card ₹7.5 lakh; complete server about ₹13–17 lakh 350 W
B — 2 × A100 80 GB (PCIe) server 100–120B mixture-of-experts models; several workloads on one node Cards ₹10–11.5 lakh each; complete server about ₹26–33 lakh 300 W per GPU
C — 8 × H100 80 GB (SXM) server 235B–400B-class models; an organisation-wide platform ₹3.5–5 crore for the complete server Up to 700 W per GPU; about 10 kW for a DGX H100-class system

Build A is where most first deployments should start. A 32 billion parameter model at 8-bit handles document question-answering, summarisation and drafting well enough for internal use, and the server fits in an ordinary rack. Build C is a data-centre purchase: at around 10 kW it draws more power than many office server rooms were built to cool, which usually means colocation.

Rent before you buy: GPU capacity inside India

On-premise is not the only way to keep data in the country. Indian GPU clouds and the government’s IndiaAI compute portal rent the same hardware by the hour from data centres in India, which turns a capital purchase into a monthly bill and lets you prove utilisation before you commit.

Per-GPU hourly rates for in-country capacity, September 2026 (before GST)

Provider GPU Rate per GPU-hour Roughly per month (730 hours)
IndiaAI Compute Portal (list rate) H100 SXM ₹153 ₹1.12 lakh
IndiaAI Compute Portal (list rate) L40S ₹67.5 ₹49,000
E2E Networks (3-month commitment) H100 ₹155.90 ₹1.14 lakh
E2E Networks (on demand) H100 ₹255.55 ₹1.87 lakh
E2E Networks (on demand) L40S ₹102 ₹74,000
Yotta Shakti Cloud (on demand) H100 SXM ₹356 ₹2.6 lakh
Yotta Shakti Cloud (on demand) L40S ₹137 ₹1 lakh

Two things stand out. First, at IndiaAI’s list rate an L40S costs about ₹49,000 a month — almost exactly what owning one costs once you spread the purchase over three years and add power, before any subsidy. The government says more than 38,000 GPUs have been onboarded at a subsidised rate of around ₹65 an hour for eligible users. Second, rented capacity is the right place to run a proof of concept: you learn your real token volume and the model size you need before you sign a purchase order for hardware you may have sized wrongly.

On-premise versus API: where break-even lands

Hosted frontier models are now cheap per token. As of September 2026, OpenAI’s GPT-6 Sol and Anthropic’s Claude Sonnet 5 both list at $2 per million input tokens and $10 per million output tokens, and Google’s Gemini 3.5 Flash at $1.50 and $9. Enterprise workloads that answer questions from documents read far more than they write, so assume three input tokens for every output token. That gives a blended price of about $4 per million tokens for the first two, or roughly ₹352 at ₹88 to the dollar.

The on-premise side has to include everything, not just the card. Our worked example spreads the purchase over three years, charges power at ₹10 per kWh with cooling overhead (a power usage effectiveness of 1.5), and costs the people at a share of a senior machine learning engineer’s salary Glassdoor India puts the average at ₹18 lakh a year.

All-in monthly cost and break-even volume against API pricing

Build All-in cost per year Per month Break-even vs ₹352 per million tokens Break-even vs Gemini 3.5 Flash (₹297)
A — 1 × L40S ₹10.4 lakh (₹5 lakh hardware, ₹0.9 lakh power, ₹4.5 lakh for a quarter of an engineer) ₹87,000 About 250 million tokens a month (8 million a day) About 290 million a month
B — 2 × A100 ₹20.3 lakh (₹10 lakh hardware, ₹1.3 lakh power, ₹9 lakh for half an engineer) ₹1.7 lakh About 480 million a month (16 million a day) About 570 million a month
C — 8 × H100 ₹1.82 crore (₹1.42 crore hardware, ₹13 lakh power, ₹27 lakh for one and a half engineers) ₹15.2 lakh About 4.3 billion a month (144 million a day) About 5.1 billion a month

Three honest caveats. The comparison is at the level of the workload, not the model: a 32 billion parameter open-weight model is not a frontier model on every task, so the break-even only counts if the open model passes your own evaluation on your own documents. The API side buys you model upgrades for free; the on-premise side buys you fixed cost, data that never leaves your network, and no per-token surprises when usage grows. And whether one node can serve your volume at the latency your users will accept is a load-testing question — which is exactly what rented GPUs are for.

Run the break-even on your own numbers

Send us your expected monthly token volume and the two or three workloads you have in mind. We will tell you which build fits, whether renting in India is the better first step, and where your break-even lands.

Talk to an engineer →

What the DPDP Act requires once the model is yours

Running the model yourself does not take you out of the Digital Personal Data Protection Act — it makes you responsible for more of it. The DPDP Rules were notified on 13 November 2025. Some provisions took effect on publication, the consent-manager provisions follow a year later, and most of the operational duties apply eighteen months after notification, which by the Rules’ own clock is May 2027. Build to them now; a deployment designed this year will still be running then.

DPDP obligations and what they mean for an on-premise LLM

Obligation Where it comes from What it means for your deployment
Reasonable security safeguards Section 8 of the Act Encryption, access control and logging cover the vector index, the prompt and response logs, and any fine-tuning data — not just the source systems. Failure here carries the highest penalty in the Act: up to ₹250 crore.
Personal data breach notification Section 8(6) of the Act; Rule 7 Notify the Data Protection Board and each affected person without delay, then send the Board a detailed report within 72 hours. Your logs must tell you whose data was in a leaked prompt or retrieved passage.
Erasure when the purpose is served or consent is withdrawn Section 8 of the Act You need a way to remove one person’s documents from the retrieval index. Personal data baked into fine-tuned weights is hard to erase — prefer retrieval over fine-tuning wherever personal data is involved.
Duties of a Significant Data Fiduciary Section 10 of the Act; Rule 13 A Data Protection Officer based in India, an independent data auditor, and a data protection impact assessment and audit every twelve months, including due diligence on algorithmic software you deploy. An LLM is algorithmic software.
Data that must stay in India Rule 13 Personal data the government restricts on a committee’s recommendation cannot leave the country. On-premise and in-country cloud meet this by design.

PIB’s summary sets out the rest of the penalty schedule: up to ₹200 crore for failing to notify a breach or for failures around children’s data, and up to ₹50 crore for other violations. The practical point for an LLM project is that the prompt log is personal data. Decide on day one how long you keep it, who can read it and how you delete from it.

The costs nobody quotes for: power, cooling, space and people

Power is the easiest to underestimate because it arrives as a separate bill. NVIDIA rates the L40S at 350 W, the A100 80 GB at 300 W in its PCIe form, and the H100 at up to 700 W in its SXM form; a complete eight-GPU H100 system draws around 10 kW at peak. Every kilowatt of computing needs cooling on top, which is why the worked example multiplies power by 1.5.

Space follows from power. A single-GPU server sits in an existing rack. An eight-GPU system usually needs a colocation cage with the power density to match, and colocation is quoted per kilowatt and per rack by each data centre — get the quote before you get the purchase order approved.

People are the line that decides whether the deployment is still healthy in year two. Someone has to own the GPU drivers, the serving stack, model upgrades, the retrieval index, capacity planning and the evaluation set that tells you a new model is actually better. For a single node that is a fraction of one engineer; for an eight-GPU platform it is more than one. Name the owner and the backup before the hardware arrives.

Procurement takes longer than the build. One US reseller quotes six to eight weeks for a DGX H100 system, and stock of cards like the L40S is patchy in India. Run the proof of concept on rented capacity while the hardware is on order.

When on-premise is the wrong answer

On-premise is a good answer to a specific problem, and a poor answer to several others. It is usually the wrong choice when:

The middle path works for many Indian enterprises: an open-weight model on GPU capacity rented inside India for the workloads that touch personal data, and a hosted frontier model, with a data-processing agreement, for tasks that do not. It keeps the data where the DPDP Act wants it without buying a server before you know you need one.

How we run this for clients

We start with a 4-week AI proof of concept on rented GPUs inside India, on your documents, with an evaluation set agreed in week one. By week four you know the model size you need, the token volume you will actually generate, the latency your users will accept, and whether the open-weight model is good enough. Only then does anyone size hardware.

If the answer is on-premise, we deploy it in your data centre or your private cloud and design the private AI infrastructure around it — serving stack, retrieval layer, access control, logging built for the DPDP obligations above, and monitoring. We then connect it to the systems it has to read from, and we stay on to run it if you want us to. Built fully custom, or accelerated by our own LLM and inference stack where it speeds delivery — the choice is yours. See how this fits into our wider enterprise AI solutions and private LLM development work.

Bring your token volume and your constraints

A 30-minute call with an engineer: your workloads, your data rules, your volume. You leave with a recommended build, a rental-first option and a rough break-even.

Book a scoping call →

Frequently asked questions

What is an on-premise LLM?

An on-premise LLM is a large language model that runs on servers you own or control — in your data centre or a private cloud — rather than through a public API. Prompts, documents and responses never leave your network, and you pay for hardware and operations instead of per token.

What hardware do you need to run an LLM on-premise?

Enough GPU memory for the model’s weights plus working headroom. As a rule of thumb, allow about 1 GB of GPU memory per billion parameters at 8-bit. A single 48 GB L40S runs 20–32 billion parameter models; 100 billion-plus models need 80 GB cards such as the A100 or H100, often two or more.

How much does it cost to run an LLM on-premise in India?

Indicatively ₹13–17 lakh for a single L40S server and ₹3.5–5 crore for an eight-GPU H100 server, as of September 2026. All-in, including power and a share of an engineer, a single-GPU deployment costs about ₹87,000 a month over three years. Renting an L40S inside India lists at ₹67.5–₹137 an hour.

Is on-premise better than a cloud API for an LLM?

Only above a steady volume, or where data cannot leave your control. Against API pricing of about ₹352 per million tokens, a single-GPU server breaks even at roughly 250 million tokens a month. Below that, an API or an in-country GPU cloud is cheaper and you keep access to model upgrades.

Can you run Claude, Gemini or OpenAI models on-premise?

The frontier Claude, Gemini and GPT models are offered as hosted services, not downloadable weights. OpenAI’s gpt-oss models and Google’s Gemma models are open-weight and can be self-hosted, as can Llama, Qwen and Mistral models, subject to each licence.

Does the DPDP Act require on-premise deployment?

No. The Act requires reasonable security safeguards, breach notification and erasure, wherever the data sits. Significant Data Fiduciaries can be required to keep specified personal data in India. On-premise and in-country cloud both satisfy that; so does a hosted service with an Indian region, if its terms allow.

How do you build RAG on an on-premise LLM?

Run an embedding model and a vector database alongside the LLM, index your documents with their access permissions attached, and retrieve only what the asking user is allowed to see. Keep the index, the prompt logs and the model on the same private network, and test answers against a fixed evaluation set before each change.

Y

yatin

Enterprise AI team at AIVeda.

← Previous

How Conversational AI Voice Bots Actually Work (Human-Like, 10+ Languages)

Next →

Data Engineering Consulting: Scope, Cost and Ownership