THE FULL GUIDE / EP 03 · ASTER × CHATGPT
Three users, a $3,000 AI bill? How to protect your AI endpoint
Your API key can be perfectly safe on the server and you can still get a bill you can’t pay. If your AI route has no login, no rate limit and no spending cap, anyone can use it on your card. Here’s how to close the pump.

Commented BILL? Here’s everything we promised: the checklist, the code, and why each step matters.
Jump to a story
- The checklist
- 01Three users, a $3,000 bill
- 02Why hiding the key isn’t enough
- 03Require login before the AI route runs
- 04Rate limit per user and per IP
- 05Using Express instead? express-rate-limit
- 06Give every user a daily budget
- 07Set spend limits and alerts in the provider dashboard
- 08How much each layer protects you
- FAQ
- Source notebook
Three users, a $3,000 bill
In the episode, Aster asks ChatGPT a simple question: why is his AI bill $3,000 when only three people use his app? He did the key part right. OPENAI_API_KEY sat on the server, in a route handler, exactly where it belongs.
The problem was the route itself. His app had a /api/chat endpoint that took any message and sent it to OpenAI. No login, no rate limit, no spending cap. Anyone who opened DevTools, saw the request and copied it could send their own prompts through his app, as many as they liked, on his card. Bots that scan for open AI endpoints don’t need to read your code, they just need the URL.
ChatGPT’s verdict in the episode: “The key is safe. The pump is wide open.” It’s a gas station that left the pump running with nobody at the counter. The fuel is locked away, but anyone can drive up and fill their tank.
# What a stranger can do with an open AI route. No key needed, yours is used.
curl -X POST 'https://your-app.com/api/chat' \
-H 'content-type: application/json' \
-d '{"message": "Write a 5,000-word essay about…"}'
# Put it in a loop and your bill grows every second.The request is visible in the Network tab of your own site. Hiding the URL doesn’t help; your frontend has to call it.
THE BUILDER’S TAKEAWAYKeeping the key on the server stops people stealing the key. It doesn’t stop them using it through your route.
Why hiding the key isn’t enough
Moving the key to the server, which is what the EP01 guide on protecting API keys is about, is step one. It means nobody can copy the key and use it from their own machine. But your server route is now a proxy for that key: it takes a request from the internet and turns it into a paid OpenAI call. If that route accepts requests from anyone, the effect on your bill is the same as publishing the key.
OWASP has a name for this: API4:2023 Unrestricted Resource Consumption. It covers APIs with no limits on how often or how much a client can use them, including APIs that call paid third-party services per request. Its prevention list reads like this episode’s fix: rate limit clients, set maximum sizes for inputs, limit how often one client can run an operation, and “configure spending limits for all service providers/API integrations”, with billing alerts where limits aren’t possible.
OpenAI’s own safety best practices say the same from the other side: users should generally need to register and log in, and limiting input length and output tokens helps reduce misuse.
It happens to real teams
- METR (2026): an AI evaluation organization disclosed that a dashboard a researcher ran on a personal cloud server had authentication that silently “failed open”. An attacker got the agent to reveal its model-provider key and used about $600,000 of credits over three weeks. At the time, there was no way to put a spending limit on that key.
- Gemini, April 2026: a developer reported a €54,000+ Gemini bill in 13 hours from an unrestricted Firebase browser key. Their budget alert and anomaly alert both fired hours late.
- A vibe-coded SaaS, March 2025: an app built entirely with an AI editor was attacked days after launch. Keys were maxed out, there was no auth, and the founder shut it down within a week.
People also report smaller versions of the same story, like a few thousand dollars of OpenAI calls they didn’t make. Those are anecdotes, but they follow the same pattern: something that spends money was reachable without a limit.
THE BUILDER’S TAKEAWAYAn AI route with no auth is a public key with a nicer URL. Treat every route that spends money like it holds the key.
Require login before the AI route runs
The first check in your route is the session. No session, no AI call, return 401. Then cap what a signed-in user can send and get back. Here’s a Next.js App Router route handler. getSession() stands in for your auth library’s server-side session check:
import 'server-only'
import OpenAI from 'openai'
// No NEXT_PUBLIC_ prefix: this value never reaches the browser bundle.
export const openai = new OpenAI({ apiKey: process.env.OPENAI_API_KEY })import { NextResponse } from 'next/server'
import { createHash } from 'node:crypto'
import { getSession } from '@/lib/session' // your auth library's server-side session check
import { openai } from '@/lib/openai'
const MAX_INPUT_CHARS = 4_000
const MAX_OUTPUT_TOKENS = 800
export async function POST(req: Request) {
// 1. No session, no AI. This runs before anything that costs money.
const session = await getSession()
const userId = session?.user?.id
if (!userId) {
return NextResponse.json({ error: 'Sign in to use chat' }, { status: 401 })
}
// 2. Validate and cap the input
const body = await req.json().catch(() => null)
const message = typeof body?.message === 'string' ? body.message.trim() : ''
if (!message || message.length > MAX_INPUT_CHARS) {
return NextResponse.json({ error: 'Message is empty or too long' }, { status: 400 })
}
// 3. Server picks the model and caps the output. The client can't change either.
const response = await openai.responses.create({
model: process.env.OPENAI_MODEL!,
instructions: 'You are the support assistant for this app. Keep answers short.',
input: message,
max_output_tokens: MAX_OUTPUT_TOKENS,
// A hashed, stable user id helps OpenAI detect abuse without getting the raw id
safety_identifier: createHash('sha256').update(userId).digest('hex'),
})
return NextResponse.json({ reply: response.output_text })
}Auth.js, Clerk, Supabase Auth and others all have a server-side way to read the session. Use that, never a flag sent by the browser.
Why every detail is there
- Session check first: everything after it can cost money, so it has to come before the OpenAI call, not after.
- A user id, not just “logged in”: you need it for the per-user rate limit and budget in the next steps.
MAX_INPUT_CHARS: input tokens are billed too. A 4,000-character cap stops someone pasting a novel into every request.max_output_tokens: an upper bound on what the model generates, including reasoning tokens. Without it, one prompt like “write 10,000 words” is billed in full.- Model from an env var: if the client can pick the model, it can pick the most expensive one.
// Anyone on the internet can call this
export async function POST(req: Request) {
const { message, model } = await req.json()
const response = await openai.responses.create({ model, input: message })
return Response.json({ reply: response.output_text })
}// Session first, server-chosen model, capped input and output
const session = await getSession()
if (!session?.user?.id) return Response.json({ error: 'Sign in' }, { status: 401 })
// …validate message length, then:
await openai.responses.create({
model: process.env.OPENAI_MODEL!,
input: message,
max_output_tokens: 800,
})THE BUILDER’S TAKEAWAYLogin turns “anyone on the internet” into “people you can identify, limit and ban”. Every other step depends on it.
Rate limit per user and per IP
Login stops strangers. A rate limit stops one account, or one script with a stolen session, from sending thousands of requests. On serverless hosts like Vercel, an in-memory counter resets with every new instance, so keep the counter in Redis. Upstash Ratelimit is built for this:
npm install @upstash/ratelimit @upstash/redis @vercel/functions# Server-only. No NEXT_PUBLIC_ prefix.
OPENAI_API_KEY=sk-...
OPENAI_MODEL=your-model-id
UPSTASH_REDIS_REST_URL=https://...
UPSTASH_REDIS_REST_TOKEN=...import 'server-only'
import { Ratelimit } from '@upstash/ratelimit'
import { Redis } from '@upstash/redis'
import { ipAddress } from '@vercel/functions'
const redis = Redis.fromEnv() // reads UPSTASH_REDIS_REST_URL and UPSTASH_REDIS_REST_TOKEN
// Signed-in users: 20 messages per minute each
const perUser = new Ratelimit({
redis,
limiter: Ratelimit.slidingWindow(20, '1 m'),
prefix: 'rl:chat:user',
})
// Per IP: a wider net for many accounts behind one address, or no session at all
const perIp = new Ratelimit({
redis,
limiter: Ratelimit.slidingWindow(60, '1 m'),
prefix: 'rl:chat:ip',
})
export function clientIp(req: Request) {
// On Vercel, ipAddress() reads the platform's IP header.
// Elsewhere, only trust x-forwarded-for if your host or proxy sets it.
return ipAddress(req) ?? req.headers.get('x-forwarded-for')?.split(',')[0]?.trim() ?? 'unknown'
}
// Limit by user id when there is one, and always by IP as well
export async function checkRateLimit(req: Request, userId?: string) {
const byIp = await perIp.limit(clientIp(req))
if (!byIp.success || !userId) return byIp
return perUser.limit(userId)
}Then add it to the route, right after the session check and before the OpenAI call:
import { checkRateLimit } from '@/lib/ratelimit'
// …after the session check
const limit = await checkRateLimit(req, userId)
if (!limit.success) {
return NextResponse.json(
{ error: 'Too many requests. Try again in a minute.' },
{ status: 429, headers: { 'Retry-After': String(Math.max(1, Math.ceil((limit.reset - Date.now()) / 1000))) } },
)
}limit() returns success, limit, remaining and reset (a Unix timestamp in milliseconds). If you turn on Upstash analytics, also wait for the returned pending promise before the function ends.
Pick numbers from your real usage
Start from what a real person can do. Someone typing in a chat box rarely sends more than a few messages a minute, so 20 per minute is generous. Watch your 429 responses for the first week: if real users hit them, raise the limit; if nobody comes close, you can lower it.
THE BUILDER’S TAKEAWAYA rate limit caps how fast money can leave. One account, one script or one bad loop can only spend so much per minute.
Using Express instead? express-rate-limit
On an Express server, express-rate-limit does the same job as middleware. Put auth first, then the limiter, then the handler:
npm install express-rate-limitimport express from 'express'
import { rateLimit, ipKeyGenerator } from 'express-rate-limit'
import OpenAI from 'openai'
const app = express()
const openai = new OpenAI({ apiKey: process.env.OPENAI_API_KEY })
// Only if you run behind exactly one proxy or load balancer, so req.ip is the real client
app.set('trust proxy', 1)
app.use(express.json({ limit: '16kb' })) // cap the request body size
// req.user is set by your session middleware (Passport, express-session, your auth SDK…)
function requireAuth(req, res, next) {
if (!req.user) return res.status(401).json({ error: 'Sign in to use chat' })
next()
}
const chatLimiter = rateLimit({
windowMs: 60 * 1000, // 1 minute
limit: 20, // 20 requests per user per minute
standardHeaders: 'draft-8',
legacyHeaders: false,
// Per user when signed in, per IP otherwise (ipKeyGenerator handles IPv6 ranges)
keyGenerator: (req) => req.user?.id ?? ipKeyGenerator(req.ip),
// With more than one server instance, add a shared store such as Redis
})
app.post('/api/chat', requireAuth, chatLimiter, async (req, res) => {
const message = typeof req.body?.message === 'string' ? req.body.message.trim() : ''
if (!message || message.length > 4000) return res.status(400).json({ error: 'Message is empty or too long' })
const response = await openai.responses.create({
model: process.env.OPENAI_MODEL,
input: message,
max_output_tokens: 800,
})
res.json({ reply: response.output_text })
})The default store is in memory, so each server instance keeps its own count. Add a Redis or Memcached store if you run more than one instance.
THE BUILDER’S TAKEAWAYSame idea, different framework: identify the caller, then cap how often they can spend your money.
Give every user a daily budget
A rate limit controls speed. A budget controls the total. 20 messages a minute is fine for a person, but a script can run at exactly that speed all day. Count the tokens each user spends per day, and stop at a limit you can afford. OpenAI returns the usage on every response, so you record real numbers, not guesses:
import 'server-only'
import { Redis } from '@upstash/redis'
const redis = Redis.fromEnv()
const DAILY_TOKENS_PER_USER = 50_000 // pick a number you can afford × your number of users
// One counter per user per UTC day
const key = (userId: string) =>
`budget:chat:${userId}:${new Date().toISOString().slice(0, 10)}`
export async function hasBudget(userId: string) {
const used = (await redis.get<number>(key(userId))) ?? 0
return used < DAILY_TOKENS_PER_USER
}
export async function recordUsage(userId: string, tokens: number) {
const k = key(userId)
await redis.incrby(k, tokens)
await redis.expire(k, 60 * 60 * 48) // clean up after two days
}import { hasBudget, recordUsage } from '@/lib/budget'
// …after the rate limit, before the OpenAI call
if (!(await hasBudget(userId))) {
return NextResponse.json({ error: 'You’ve reached today’s chat limit.' }, { status: 429 })
}
const response = await openai.responses.create({ /* …as in step 1 */ })
// Record what this request actually used (input + output tokens)
await recordUsage(userId, response.usage?.total_tokens ?? 0)Prefer your database? The same counter works as one row per user per day:
create table ai_usage (
user_id text not null,
day date not null default current_date,
tokens bigint not null default 0,
primary key (user_id, day)
);
-- after each response: add the tokens it used
insert into ai_usage (user_id, day, tokens)
values ($1, current_date, $2)
on conflict (user_id, day)
do update set tokens = ai_usage.tokens + excluded.tokens;Tokens or dollars?
Tokens are what the API reports, so they’re the simplest thing to count. If you use several models, multiply input and output tokens by each model’s price from OpenAI’s pricing page and store cents instead. Either way, the check happens before the call and the record happens after it, so a few requests already in flight can go slightly over. The rate limit keeps that overshoot small.
THE BUILDER’S TAKEAWAYRate limits stop bursts. Budgets stop the slow, steady drain that runs all night at exactly your limit.
Set spend limits and alerts in the provider dashboard
Your code is the first net. The provider’s billing controls are the last one, for the day your code has a bug. Here’s what we could confirm in the official docs as of October 2026.
OpenAI
- Hard spend limits at two levels: for the whole organization (Settings → Organization limits → Spend) and per project (Project settings → Limits → Spend). When a hard limit is reached, requests return
429withorganization_spend_limit_exceededorproject_spend_limit_exceeded. Enforcement isn’t instant, so recorded spend can go slightly over. - Spend alerts email you when usage passes an amount. They don’t stop traffic, and they keep working alongside a hard limit, so set alerts below the limit (for example at 50% and 80%).
- Separate projects for staging and production, each with its own keys and its own rate and spend limits.
- Prepaid billing: new API accounts buy credits up front and usage stops when the balance runs out. Auto-recharge is on by default when you set it up. You can turn it off, or keep it with an optional monthly recharge limit.
- Usage tiers come with a monthly usage limit for your organization that rises as you spend more. It’s a platform ceiling, not a budget you chose, so set your own lower limit.
Google Gemini API and Vertex AI
- Project spend caps in Google AI Studio (announced March 2026): a monthly dollar limit for Gemini API spend per project. Google marks it experimental and notes about a 10-minute billing delay, and you pay for overages in that window.
- Usage tier caps: each Gemini API tier has a maximum monthly spend across the billing account (Tier 1 is $250). When it’s reached, service pauses for all linked projects until the next month.
- Cloud Billing spend cap budgets (Preview) can block new usage of the Gemini API and Vertex AI (now Gemini Enterprise Agent Platform) in a project once a budget is passed, until you lift the cap. Google notes enforcement isn’t instant and overages are billed.
- Restrict your Google API keys. Keys that were meant for Maps or Firebase could start working with Gemini once the API was enabled on the project. Limit each key to the APIs it needs.
THE BUILDER’S TAKEAWAYSet the hard limit to the most you could pay without panicking. Then set alerts well below it, so you hear about trouble first.
- OpenAI: spend limits
- OpenAI: production best practices
- OpenAI Help: prepaid API billing
- OpenAI: rate limits
- Gemini API: billing and spend caps
- Google: more control over Gemini API costs (Mar 2026)
- Google Cloud Billing: spend cap budgets
- Truffle Security: Google API keys weren’t secrets, then Gemini changed the rules
- Google AI Developers Forum: €54k billing spike in 13 hours
How much each layer protects you
| Setup | Who can spend your money | Worst case on a bad night | What stops it |
|---|---|---|---|
| No protection | Anyone who finds the URL, including bots | Unlimited, until you notice or the card fails | Nothing |
| Login only | Anyone who signs up, as fast as they can script it | One free account can still run a loop all night | You, when you spot it and ban the account |
| Login + rate limit | Signed-in users, at a human speed | Users × your per-minute limit × hours awake | The rate limit slows it, but doesn’t end it |
| All three (login, rate limit, spend caps) | Signed-in users, at a human speed, up to a daily budget | Each user’s daily budget, and never more than your provider hard limit | Per-user budget first, provider hard limit last |
THE BUILDER’S TAKEAWAYEach layer covers a hole in the one before it. Only all three together put a ceiling on the bill.
QUESTIONS PEOPLE ASK
FAQ
Why is my OpenAI API bill so high when I have few users?
Usually because something other than your users is calling your AI route. If /api/chat (or any route that calls OpenAI) works without a login, anyone who finds it can send requests on your key. Check the OpenAI usage dashboard for traffic that doesn’t match your real users, then add auth, a rate limit and spend limits.
Is my OpenAI API key safe if it’s only on the server?
The key itself is safe from being copied, but your route can still be abused. A server route that forwards any request to OpenAI works like a public key. Require a login, rate limit each user and cap spend.
Can I set a hard spending limit on the OpenAI API?
Yes. OpenAI supports hard spend limits for the organization and for each project, set in the dashboard’s limits settings. When a limit is reached, requests return a 429 error until you raise it or the next monthly cycle starts. Enforcement isn’t instant, so spend can go slightly over.
What’s the difference between an OpenAI spend alert and a spend limit?
An alert sends a notification when spend passes an amount, and traffic keeps flowing. A hard spend limit stops requests. Use both: alerts below the limit so you hear about trouble before the limit cuts traffic.
How do I rate limit a Next.js API route?
Use a shared store, because serverless instances don’t share memory. With Upstash Ratelimit, create a Ratelimit with Ratelimit.slidingWindow(20, '1 m'), call limit(userId) at the start of the route, and return 429 when success is false. Fall back to the client IP when there’s no user.
Should I rate limit by user or by IP?
Both. Per-user limits are fair and can’t be dodged by switching networks. Per-IP limits catch many accounts from one address and traffic without a session. Behind a proxy, make sure you read the real client IP from a header your host controls.
How do I limit how much each user can spend on AI?
Keep a per-user daily counter in Redis or your database. Check it before each AI call, and after the call add the tokens from the response’s usage field. When the counter passes your limit, return a friendly “daily limit reached” message.
Does Google offer spend caps for the Gemini API?
Yes, as of 2026. Google AI Studio has project spend caps (experimental, with about a 10-minute delay), each usage tier has a monthly spend maximum, and Cloud Billing has spend cap budgets in Preview for the Gemini API and Vertex AI. None of them are instant, so keep limits in your own code too.
THE BIGGER PICTURE
The takeaway for builders.
Keeping your AI key on the server is the right start, but a route that spends money needs a lock of its own. Require login before the AI runs, rate limit per user and IP, give each user a daily budget, and set a hard spend limit with your provider for the night something slips through. A free Aster repository audit can take a second look at the code side, like where your keys end up.
The source notebook 21 links
Go a little deeper. These are the sources linked in this story.
- 01OWASP API Security Top 10: API4:2023 Unrestricted Resource Consumptionapi-security.owasp.org
- 02OpenAI: safety best practicesdevelopers.openai.com
- 03OpenAI: best practices for API key safetyhelp.openai.com
- 04METR: security update (Aug 2026)metr.org
- 05Google AI Developers Forum: €54k billing spike in 13 hoursdiscuss.ai.google.dev
- 06Pivot to AI: a vibe-coded SaaS under attack (Mar 2025)pivot-to-ai.com
- 07OpenAI API reference: create a model responsedevelopers.openai.com
- 08Auth.js: protecting resourcesauthjs.dev
- 09Upstash: Ratelimit getting startedupstash.com
- 10Upstash: Ratelimit methodsupstash.com
- 11Vercel: @vercel/functions API referencevercel.com
- 12express-rate-limit READMEgithub.com
- 13express-rate-limit: configurationexpress-rate-limit.mintlify.app
- 14OpenAI: rate limitsdevelopers.openai.com
- 15OpenAI: spend limitsdevelopers.openai.com
- 16OpenAI: production best practicesdevelopers.openai.com
- 17OpenAI Help: prepaid API billinghelp.openai.com
- 18Gemini API: billing and spend capsai.google.dev
- 19Google: more control over Gemini API costs (Mar 2026)blog.google
- 20Google Cloud Billing: spend cap budgetsdocs.cloud.google.com
- 21Truffle Security: Google API keys weren’t secrets, then Gemini changed the rulestrufflesecurity.com
PUT THE CONTEXT TO WORK
A second look.
Before you ship.
Connect your GitHub repository and see what needs attention. Your first audit is free.
Start my free audit

