A document-analysis feature can accept one upload and create many downstream operations. It may extract pages, generate embeddings, retrieve related records, and run several model calls before returning a result. Counting incoming requests alone does not describe the resources the feature consumes.

For a business offering AI functionality, usage controls affect both expenditure and service availability. One customer's large job can occupy capacity that other customers need, even when every request is authenticated.

Account for the work behind each request

OWASP includes unbounded consumption among its risks for LLM applications. The category covers uncontrolled resource use, including excessive inference and related operational costs. Its guidance discusses limits, monitoring, and controls across the processing path.

An illustrative contract-review service might charge customers per document. A two-page agreement and a scanned archive of hundreds of pages can have different processing requirements. Internally, a document count alone cannot distinguish them. Page count, extracted text size, model usage, elapsed processing time, and concurrent jobs can provide additional measures.

Retries can change the workload

A timeout does not always mean the underlying task stopped. If the client resubmits while the original job continues, the service may perform the same expensive work twice. Automatic retries between internal services can add further duplication.

Assigning a stable job identity lets the application connect retries to existing work. Define which failures permit retry, how many attempts are allowed, and when the job becomes terminal. A cancellation request also needs a documented meaning: whether it stops queued work, signals running workers, or simply stops the user interface from waiting.

Test limits without generating a production incident

Use a dedicated environment with synthetic documents and a small explicit spending ceiling. Submit a normal document, a document near the supported size limit, and several simultaneous jobs from one test tenant. Then submit work from a second tenant and observe whether it receives its expected share of capacity.

Introduce a controlled downstream timeout and inspect the number of actual processing attempts. Compare the front-end status, queue state, provider usage records, and internal job log. These records can reveal work that continues after the application reports cancellation or failure.

Tie the result to a product decision

The product owner can define an allowed workload for each account and an understandable response when that boundary is reached. A queue, a deferred job, and a rejected request are different service behaviors. The customer-facing message should match the action the system actually takes.

An assessment can document the maximum observed job expansion, the enforcement point for each limit, and the response under contention. Those findings support capacity planning and abuse controls without implying that a particular token limit guarantees availability. Recheck the behavior when adding tools, longer context, batch processing, or automatic retries because each can change the work triggered by one request.

Sources

OWASP: Denial of Service. The contract-review service and workload tests are illustrative.

Back to the blogExplore Ceron