Customer case study

Impact in cost and key metrics for a production LLM inference provider

Representative inference deployment · 2026 field analysis

Last updated 21 August 2026

On a production LLM inference workload, Feniria cut monthly GPU cost by 50% and p99 time to first token by 48%, with no hardware added and no code changed. Throughput rose 9% and end-to-end latency fell 37% on the same GPUs.

−27%
Energy for the same workload
less kWh for the same work
−48%
p99 time to first token
faster first response
+9%
Inference throughput
more tokens per second
−37%
End-to-end latency
shorter request completion

Cost and productivity impact

50% lower monthly cost. Same workload.

Baseline: single g6e.48xlarge on AWS. Feniria found unused capacity and optimized the deployment. No code or model changes.

Baseline

$21,996

Total monthly cost

Feniria optimized

$10,998

Total monthly cost · 50% lower

Total monthly cost
Baseline $21,996
Feniria $10,998

50% lower cost

Waste
Baseline 89%
Feniria 20%

Idle capacity, over-provisioning, misconfiguration → minimal and well-managed

Productive
Baseline 11%
Feniria 80%

More output, better usage, right-sized

Productivity impact

Utilisation dashboards looked healthy the whole time. These are the metrics that showed the gap between reported utilisation and actual productive work, and closed it.

p99 time to first token
Change −48%
Impact Faster first response
Inference throughput
Change +9%
Impact More tokens per second
End-to-end latency
Change −37%
Impact Shorter request completion
Core idle time (proxy via core-hours)
Change −19pp
Impact Less unproductive capacity
Total energy (active + idle) for the same workload
Change −27%
Impact Less kWh for the same inference load

Core idle time is calculated as (1 − SM Active), averaged across the full measurement window, including idle and load periods.

In real-world terms

More compute value. Less energy wasted.

For a representative 100-GPU fleet running the same workload, this deployment's measured +19% capacity gain and −27% energy reduction translate into real cloud compute value and energy savings every month, no new hardware required.

Cloud compute value

~€85,000 /month

Equivalent to running 19 extra GPUs for a full month, on the fleet you already have.

Energy saved

~32 MWh /month

27% less energy for the same workload.

CO₂ avoided

~9.4 t CO₂/month

≈112 t CO₂/year (2024 EU average). Equal to what ~5,100 mature trees absorb in a year.

  1. Illustrative example for a representative 100-GPU H100 fleet running 24/7. Cloud compute value uses this deployment's measured +19% effective capacity gain. Energy and CO₂ use the measured −27% energy reduction for the same workload.
  2. GPU cost: €6.20/GPU-hr (~$6.88 AWS on-demand). $6.88 is AWS p5.48xlarge (8× H100 at $55.04/h) divided per GPU. Operating hours: 720/month (30 days × 24 h).
  3. Power draw: 1.67 kW per GPU (700 W TDP × 1.7 overhead × 1.4 PUE). Baseline fleet: 100 × 1.67 kW = 167 kW, or 167 kW × 720 h = 120.24 MWh/month. Energy saved at −27%: 0.27 × 120.24 ≈ 32.46 MWh/month, shown as ~32 MWh.
  4. Grid carbon intensity: 288 kg CO₂/MWh, 2024 EU aggregate consumption mix (Fraunhofer ISE). Ember's 2024 EU generation intensity is ~213 g CO₂/kWh; 288 is the more conservative consumption-mix factor. Avoided CO₂: 32.46 MWh × 288 kg/MWh ≈ 9.35 t/month, shown as ~9.4 t (≈112 t/year).
  5. Tree equivalence: ~22 kg CO₂ per mature tree per year (US EPA / Arbor Day Foundation). Annual avoided CO₂ of 112,200 kg ÷ 22 kg/tree ≈ 5,100 trees absorbing that amount in one year.

The problem

Utilisation is not productivity

AI infrastructure tends to scale faster than teams can use it well. Most GPU fleets run well below productive capacity.

Standard dashboards report high utilisation while productive compute stays low. That gap leads to hardware purchases, energy draw, and carbon load that could have been avoided.

Representative field result

What we found and what we changed

What we found

The deployment's standard dashboard reported high GPU utilisation. Hardware-level telemetry from Feniria told a different story. During normal traffic periods, SM Active (the fraction of streaming multiprocessors doing productive compute work) was in single digits, meaning roughly 90% of processing capacity was idle at those moments. Across the full measurement window, including peak load, an average of 60% of available core hours were idle or unproductive, peaking at 82%.

What we changed

Feniria applied that combined picture, utilisation, throughput, latency, and hardware-level metrics read together, to optimize the LLM serving engine and scale the deployment's concurrency. The objective throughout was low-level optimization: getting more productive work out of the GPUs already running, not adding hardware.

What teams usually want to know

Questions about this case study

What did Feniria change on this production deployment?

Feniria optimized the LLM serving engine and concurrency using hardware-level telemetry, not extra GPUs. No code or model changes.

How was the 50% cost reduction measured?

Baseline was a single AWS g6e.48xlarge at $21,996 per month. After Feniria found unused capacity and tuned the deployment, the same workload cost $10,998 per month.

Did utilization dashboards already look healthy?

Yes. Standard dashboards reported high GPU utilization while SM Active sat in single digits during normal traffic. Across the full measurement window, about 60% of core hours were idle or unproductive.

Start fixing

We find the waste. We fix it.

Join our early adopters, with limited slots for teams ready to ship more with the GPUs they already have. Piloting now with design partners, on cloud and on-prem.

Request early access