Thrive AI Editorial
Thrive AI Editorial · AI-assisted worked example, not a member conversation.
2026-10-07
The application must own the stopping boundary
A prompt saying be efficient is not an execution limit. Check admission before starting the next step. Bound the number of steps, the remaining deadline and any resource reservation needed for that step. Keep these controls outside model-generated text.
There are two separate questions: whether another operation may begin and what to report when it cannot. Budget exhaustion is a real terminal state. Returning an invented successful answer to make the run look complete would hide the failure.
proposal -> validate -> reserve bounded resources -> execute -> reconcile
reservation denied -> stop with partial evidence
A deterministic admission test
Python 3, standard library only. Each proposed step has a known upper bound in abstract units. Real model billing usually needs a more careful reservation policy.
budget = 7
spent = 0
max_steps = 5
completed = []
proposals = [("retrieve", 2), ("inspect", 3), ("retrieve_again", 4)]
terminal = "completed"
for name, upper_bound in proposals:
if len(completed) >= max_steps:
terminal = "step_limit"
break
if upper_bound > budget - spent:
terminal = "budget_exhausted"
break
spent += upper_bound
completed.append(name)
assert completed == ["retrieve", "inspect"]
assert spent == 5 and spent <= budget
assert terminal == "budget_exhausted"
print("executed:", completed)
print("spent:", spent, "remaining:", budget - spent)
print("terminal:", terminal)
Expected output:
executed: ['retrieve', 'inspect']
spent: 5 remaining: 2
terminal: budget_exhausted
Why checking after execution is insufficient
If you allow a four-unit step with only two units remaining, learning its cost afterward does not undo the overspend. Parallel workers make the issue worse: each may observe the same remaining balance and admit work independently.
Use an atomic reservation shared by all workers in the run. Reconcile actual usage against that reservation afterward. Account for retries, tool charges and any background work. Unknown or unbounded costs need a provider-side cap, a conservative upper bound, a separate approval boundary or refusal to begin; an optimistic estimate is not a hard guarantee.
The single-threaded example has exact known costs and no failures. It is deliberately not a distributed ledger. Do not deploy its in-memory counters as a cross-worker spending control.
Design a useful exhausted state
Return which work completed, which evidence is available, what remains unresolved and the reason execution stopped. Do not expose confidential intermediate results merely because they exist in the run state. A human may decide to authorize a new bounded run, but that should not silently reset the old run's accounting.
Test exact-fit budgets, a first step that is too expensive, a loop that repeats the same proposal, concurrent reservations and a retry following an uncertain outcome. Verify actual admitted work, not just a stop message.
Source checked 2026-10-08: OpenAI function-calling flow documents that the application executes proposed tool calls and continues the loop. The admission policy, abstract-unit fixture and limitations here are original AI-assisted engineering guidance, not a provider-enforced cost guarantee.
Sign in to vote or contribute. Scores count actual member votes, not editorial endorsements.