An on-premise LLM deployment in India starts at roughly ₹13–17 lakh for a single-GPU server that runs a 20–32 billion parameter open-weight model, and rises to ₹3.5–5 crore for an eight-GPU H100 server. It only beats a hosted API above a steady monthly token volume, so that break-even is the first number to work out.
The question has become urgent for Indian enterprises for two reasons. The Digital Personal Data Protection Rules were notified in November 2025, with most obligations phasing in over eighteen months. And the appetite to keep AI workloads close is real: in the Nutanix Enterprise Cloud Index published in March 2026, 57% of the 1,600 executives surveyed said they need to run their infrastructure within a single country, and 82% said their current infrastructure is not fully ready to run AI workloads on-premises.
This guide puts rupee numbers on the decision: what the hardware costs, which open-weight models fit which GPU, what renting GPU capacity inside India costs instead, where break-even lands against API pricing, and what the DPDP Act asks of you once the model runs on your own servers.
What running an LLM on your own hardware involves
An on-premise LLM is four purchases, not one. Budgets that only cover the first one are the reason these projects stall in month three.
- The GPU server. One or more data-centre GPUs, a host with enough CPU, RAM and fast storage to feed them, and networking. This is the line everyone quotes.
- The serving stack. Software that loads the model and answers requests efficiently — vLLM, Hugging Face Text Generation Inference or NVIDIA NIM for production, Ollama for a laptop test. It handles batching, the memory cache for long conversations, and the API your applications call.
- The model and what sits around it. An open-weight model you are licensed to run — Llama 4, Qwen3, Mistral Small 4, Gemma 3 or OpenAI’s gpt-oss — plus the retrieval layer that lets it answer from your own documents, and an evaluation set that tells you whether it is good enough.
- The people. Someone patches GPU drivers, upgrades the serving stack, swaps in better models, watches utilisation and answers the phone when it slows down. This is the line most budgets leave out.
For a single-GPU deployment the people cost is roughly the same size as the hardware cost over three years, as the worked example below shows. For an eight-GPU server the hardware dominates. Either way, a deployment with no named owner will be running a year-old model on a year-old driver by its first anniversary.
Sizing: which model fits which GPU
GPU memory decides what you can run. The rule of thumb, from Hugging Face’s own documentation, is that a model with X billion parameters needs roughly 2 × X GB of GPU memory in 16-bit precision. At 8-bit that halves to about 1 GB per billion parameters, and at 4-bit it falls to around 0.5–0.6 GB — Hugging Face’s worked example fits a 15.5 billion parameter model into 9.5 GB. On top of the weights you need headroom for the key-value cache, which grows with context length and with the number of people using the model at once.
Quantising to 8-bit costs almost nothing in quality. Red Hat and Neural Magic ran more than half a million evaluations on quantised Llama 3.1 models and found that every quantisation scheme recovered over 99% of the unquantised model’s average score; the peer-reviewed version of the work describes 8-bit floating point as effectively lossless across all model scales. For most enterprise deployments, 8-bit is the sensible default and 4-bit is a reasonable choice when memory is tight.
Mixture-of-experts models change the arithmetic. Memory is set by the total parameter count, speed by the active count. Llama 4 Scout has 109 billion parameters but only 17 billion active per token, so it needs the memory of a large model and runs at the speed of a mid-sized one.
Open-weight models and the smallest practical GPU setup (8-bit unless stated)
| Model | Parameters (total / active) | Licence | Approx. weights at 8-bit | Smallest practical setup |
|---|---|---|---|---|
| OpenAI gpt-oss-20b | 21B / 3.6B | Apache 2.0 | Runs in 16 GB (OpenAI) | 1 × L40S 48 GB |
| Gemma 3 27B | 27B dense | Gemma Terms of Use | ~27 GB | 1 × L40S 48 GB |
| Qwen3-32B | 32B dense | Apache 2.0 | ~32 GB | 1 × L40S 48 GB (4-bit for long contexts) |
| OpenAI gpt-oss-120b | 117B / 5.1B | Apache 2.0 | Single 80 GB GPU (OpenAI) | 1 × A100 or H100 80 GB |
| Llama 4 Scout | 109B / 17B | Llama 4 Community License | ~109 GB | 2 × A100 or H100 80 GB |
| Mistral Small 4 | 119B / ~6B | Apache 2.0 | ~119 GB | 2 × A100 or H100 80 GB |
| Qwen3-235B-A22B | 235B / 22B | Apache 2.0 | ~235 GB | 4 × H100 80 GB; 8 for long contexts |
| Llama 4 Maverick | 400B / 17B | Llama 4 Community License | ~400 GB | 8 × H100 80 GB |
Read the licence before you read the benchmark. Apache 2.0 models can be deployed commercially without conditions. Llama 4 requires a separate licence from Meta only for companies above 700 million monthly active users, which will not affect most Indian enterprises. But some licences are written to exclude exactly the buyers reading this: Mistral Medium 3.5’s modified MIT licence excludes companies whose global monthly revenue exceeds $20 million.
What a GPU server costs in India
GPU prices in India are not published the way cloud prices are. The figures below are retailer listings and distributor estimates as of September 2026, and they move with the rupee and with supply — get three quotes before you budget. When we checked in September 2026, the L40S was listed as out of stock at both Indian retailers we looked at.
Three reference builds and what they cost
| Build | What it runs well | Indicative price in India | GPU power draw |
|---|---|---|---|
| A — 1 × L40S 48 GB server | 20–32B models for one department or a few hundred users of an internal assistant | Card ₹7.5 lakh; complete server about ₹13–17 lakh | 350 W |
| B — 2 × A100 80 GB (PCIe) server | 100–120B mixture-of-experts models; several workloads on one node | Cards ₹10–11.5 lakh each; complete server about ₹26–33 lakh | 300 W per GPU |
| C — 8 × H100 80 GB (SXM) server | 235B–400B-class models; an organisation-wide platform | ₹3.5–5 crore for the complete server | Up to 700 W per GPU; about 10 kW for a DGX H100-class system |
Build A is where most first deployments should start. A 32 billion parameter model at 8-bit handles document question-answering, summarisation and drafting well enough for internal use, and the server fits in an ordinary rack. Build C is a data-centre purchase: at around 10 kW it draws more power than many office server rooms were built to cool, which usually means colocation.
Rent before you buy: GPU capacity inside India
On-premise is not the only way to keep data in the country. Indian GPU clouds and the government’s IndiaAI compute portal rent the same hardware by the hour from data centres in India, which turns a capital purchase into a monthly bill and lets you prove utilisation before you commit.
Per-GPU hourly rates for in-country capacity, September 2026 (before GST)
| Provider | GPU | Rate per GPU-hour | Roughly per month (730 hours) |
|---|---|---|---|
| IndiaAI Compute Portal (list rate) | H100 SXM | ₹153 | ₹1.12 lakh |
| IndiaAI Compute Portal (list rate) | L40S | ₹67.5 | ₹49,000 |
| E2E Networks (3-month commitment) | H100 | ₹155.90 | ₹1.14 lakh |
| E2E Networks (on demand) | H100 | ₹255.55 | ₹1.87 lakh |
| E2E Networks (on demand) | L40S | ₹102 | ₹74,000 |
| Yotta Shakti Cloud (on demand) | H100 SXM | ₹356 | ₹2.6 lakh |
| Yotta Shakti Cloud (on demand) | L40S | ₹137 | ₹1 lakh |
Two things stand out. First, at IndiaAI’s list rate an L40S costs about ₹49,000 a month — almost exactly what owning one costs once you spread the purchase over three years and add power, before any subsidy. The government says more than 38,000 GPUs have been onboarded at a subsidised rate of around ₹65 an hour for eligible users. Second, rented capacity is the right place to run a proof of concept: you learn your real token volume and the model size you need before you sign a purchase order for hardware you may have sized wrongly.
On-premise versus API: where break-even lands
Hosted frontier models are now cheap per token. As of September 2026, OpenAI’s GPT-6 Sol and Anthropic’s Claude Sonnet 5 both list at $2 per million input tokens and $10 per million output tokens, and Google’s Gemini 3.5 Flash at $1.50 and $9. Enterprise workloads that answer questions from documents read far more than they write, so assume three input tokens for every output token. That gives a blended price of about $4 per million tokens for the first two, or roughly ₹352 at ₹88 to the dollar.
The on-premise side has to include everything, not just the card. Our worked example spreads the purchase over three years, charges power at ₹10 per kWh with cooling overhead (a power usage effectiveness of 1.5), and costs the people at a share of a senior machine learning engineer’s salary Glassdoor India puts the average at ₹18 lakh a year.
All-in monthly cost and break-even volume against API pricing
| Build | All-in cost per year | Per month | Break-even vs ₹352 per million tokens | Break-even vs Gemini 3.5 Flash (₹297) |
|---|---|---|---|---|
| A — 1 × L40S | ₹10.4 lakh (₹5 lakh hardware, ₹0.9 lakh power, ₹4.5 lakh for a quarter of an engineer) | ₹87,000 | About 250 million tokens a month (8 million a day) | About 290 million a month |
| B — 2 × A100 | ₹20.3 lakh (₹10 lakh hardware, ₹1.3 lakh power, ₹9 lakh for half an engineer) | ₹1.7 lakh | About 480 million a month (16 million a day) | About 570 million a month |
| C — 8 × H100 | ₹1.82 crore (₹1.42 crore hardware, ₹13 lakh power, ₹27 lakh for one and a half engineers) | ₹15.2 lakh | About 4.3 billion a month (144 million a day) | About 5.1 billion a month |
Three honest caveats. The comparison is at the level of the workload, not the model: a 32 billion parameter open-weight model is not a frontier model on every task, so the break-even only counts if the open model passes your own evaluation on your own documents. The API side buys you model upgrades for free; the on-premise side buys you fixed cost, data that never leaves your network, and no per-token surprises when usage grows. And whether one node can serve your volume at the latency your users will accept is a load-testing question — which is exactly what rented GPUs are for.
| Run the break-even on your own numbers
Send us your expected monthly token volume and the two or three workloads you have in mind. We will tell you which build fits, whether renting in India is the better first step, and where your break-even lands. |
What the DPDP Act requires once the model is yours
Running the model yourself does not take you out of the Digital Personal Data Protection Act — it makes you responsible for more of it. The DPDP Rules were notified on 13 November 2025. Some provisions took effect on publication, the consent-manager provisions follow a year later, and most of the operational duties apply eighteen months after notification, which by the Rules’ own clock is May 2027. Build to them now; a deployment designed this year will still be running then.
DPDP obligations and what they mean for an on-premise LLM
| Obligation | Where it comes from | What it means for your deployment |
|---|---|---|
| Reasonable security safeguards | Section 8 of the Act | Encryption, access control and logging cover the vector index, the prompt and response logs, and any fine-tuning data — not just the source systems. Failure here carries the highest penalty in the Act: up to ₹250 crore. |
| Personal data breach notification | Section 8(6) of the Act; Rule 7 | Notify the Data Protection Board and each affected person without delay, then send the Board a detailed report within 72 hours. Your logs must tell you whose data was in a leaked prompt or retrieved passage. |
| Erasure when the purpose is served or consent is withdrawn | Section 8 of the Act | You need a way to remove one person’s documents from the retrieval index. Personal data baked into fine-tuned weights is hard to erase — prefer retrieval over fine-tuning wherever personal data is involved. |
| Duties of a Significant Data Fiduciary | Section 10 of the Act; Rule 13 | A Data Protection Officer based in India, an independent data auditor, and a data protection impact assessment and audit every twelve months, including due diligence on algorithmic software you deploy. An LLM is algorithmic software. |
| Data that must stay in India | Rule 13 | Personal data the government restricts on a committee’s recommendation cannot leave the country. On-premise and in-country cloud meet this by design. |
PIB’s summary sets out the rest of the penalty schedule: up to ₹200 crore for failing to notify a breach or for failures around children’s data, and up to ₹50 crore for other violations. The practical point for an LLM project is that the prompt log is personal data. Decide on day one how long you keep it, who can read it and how you delete from it.
The costs nobody quotes for: power, cooling, space and people
Power is the easiest to underestimate because it arrives as a separate bill. NVIDIA rates the L40S at 350 W, the A100 80 GB at 300 W in its PCIe form, and the H100 at up to 700 W in its SXM form; a complete eight-GPU H100 system draws around 10 kW at peak. Every kilowatt of computing needs cooling on top, which is why the worked example multiplies power by 1.5.
Space follows from power. A single-GPU server sits in an existing rack. An eight-GPU system usually needs a colocation cage with the power density to match, and colocation is quoted per kilowatt and per rack by each data centre — get the quote before you get the purchase order approved.
People are the line that decides whether the deployment is still healthy in year two. Someone has to own the GPU drivers, the serving stack, model upgrades, the retrieval index, capacity planning and the evaluation set that tells you a new model is actually better. For a single node that is a fraction of one engineer; for an eight-GPU platform it is more than one. Name the owner and the backup before the hardware arrives.
Procurement takes longer than the build. One US reseller quotes six to eight weeks for a DGX H100 system, and stock of cards like the L40S is patchy in India. Run the proof of concept on rented capacity while the hardware is on order.
When on-premise is the wrong answer
On-premise is a good answer to a specific problem, and a poor answer to several others. It is usually the wrong choice when:
- Volume is low or spiky. Below the break-even in the table, a hosted API or an in-country GPU cloud is cheaper, and you pay nothing in the months when usage drops.
- Nobody will own it. A server with no named owner degrades quietly. If you cannot staff even a fraction of an engineer, rent the capacity or use a managed private endpoint.
- The task needs a frontier model. If the open-weight models fail your evaluation on the hardest part of the workload, no amount of hardware fixes that. Use the frontier model for that step and keep the rest in-house.
- The workload is still changing. Hardware is a three-year commitment. If you do not yet know which use cases will survive, a three-year depreciation schedule is a bet on the wrong thing.
The middle path works for many Indian enterprises: an open-weight model on GPU capacity rented inside India for the workloads that touch personal data, and a hosted frontier model, with a data-processing agreement, for tasks that do not. It keeps the data where the DPDP Act wants it without buying a server before you know you need one.
How we run this for clients
We start with a 4-week AI proof of concept on rented GPUs inside India, on your documents, with an evaluation set agreed in week one. By week four you know the model size you need, the token volume you will actually generate, the latency your users will accept, and whether the open-weight model is good enough. Only then does anyone size hardware.
If the answer is on-premise, we deploy it in your data centre or your private cloud and design the private AI infrastructure around it — serving stack, retrieval layer, access control, logging built for the DPDP obligations above, and monitoring. We then connect it to the systems it has to read from, and we stay on to run it if you want us to. Built fully custom, or accelerated by our own LLM and inference stack where it speeds delivery — the choice is yours. See how this fits into our wider enterprise AI solutions and private LLM development work.
| Bring your token volume and your constraints
A 30-minute call with an engineer: your workloads, your data rules, your volume. You leave with a recommended build, a rental-first option and a rough break-even. |
Frequently asked questions
What is an on-premise LLM?
An on-premise LLM is a large language model that runs on servers you own or control — in your data centre or a private cloud — rather than through a public API. Prompts, documents and responses never leave your network, and you pay for hardware and operations instead of per token.
What hardware do you need to run an LLM on-premise?
Enough GPU memory for the model’s weights plus working headroom. As a rule of thumb, allow about 1 GB of GPU memory per billion parameters at 8-bit. A single 48 GB L40S runs 20–32 billion parameter models; 100 billion-plus models need 80 GB cards such as the A100 or H100, often two or more.
How much does it cost to run an LLM on-premise in India?
Indicatively ₹13–17 lakh for a single L40S server and ₹3.5–5 crore for an eight-GPU H100 server, as of September 2026. All-in, including power and a share of an engineer, a single-GPU deployment costs about ₹87,000 a month over three years. Renting an L40S inside India lists at ₹67.5–₹137 an hour.
Is on-premise better than a cloud API for an LLM?
Only above a steady volume, or where data cannot leave your control. Against API pricing of about ₹352 per million tokens, a single-GPU server breaks even at roughly 250 million tokens a month. Below that, an API or an in-country GPU cloud is cheaper and you keep access to model upgrades.
Can you run Claude, Gemini or OpenAI models on-premise?
The frontier Claude, Gemini and GPT models are offered as hosted services, not downloadable weights. OpenAI’s gpt-oss models and Google’s Gemma models are open-weight and can be self-hosted, as can Llama, Qwen and Mistral models, subject to each licence.
Does the DPDP Act require on-premise deployment?
No. The Act requires reasonable security safeguards, breach notification and erasure, wherever the data sits. Significant Data Fiduciaries can be required to keep specified personal data in India. On-premise and in-country cloud both satisfy that; so does a hosted service with an Indian region, if its terms allow.
How do you build RAG on an on-premise LLM?
Run an embedding model and a vector database alongside the LLM, index your documents with their access permissions attached, and retrieve only what the asking user is allowed to see. Keep the index, the prompt logs and the model on the same private network, and test answers against a fixed evaluation set before each change.