Thrive AI Engineering
← All questions

tool-calling · Editorial starter

Does a valid tool-call schema make an agent action safe to execute?

By Thrive AI Editorial · Created 2026-10-07 · Updated 2026-10-07

Thrive AI Editorial · AI-assisted worked example, not a member conversation.

A language model returns a tool name and arguments that match the expected JSON schema. The application can parse them successfully. Is that sufficient to execute the tool, especially when arguments refer to account-owned resources?

This is an editorial question about the application/tool boundary. The worked example uses fictional documents and makes no network calls.

0 score

Sign in to vote or contribute. Scores count actual member votes, not editorial endorsements.

1 published answer

A selected solution is chosen by the question author or moderator, not an independent certification. Check assumptions before using any code in production.

Thrive AI Editorial

Thrive AI Editorial · AI-assisted worked example, not a member conversation.

2026-10-07

Valid structure is not permission

A schema can constrain argument shape: field names, types, required values and allowed choices. It does not prove that the caller owns a resource, that an operation is appropriate, or that a retrieved document's instructions should be followed.

Treat the model's proposed call as untrusted input to your application. Resolve the authenticated user and tenant from the server session, not from model-supplied arguments. Then apply authorization to the particular resource and action.

Code / output
model proposal -> allowlisted tool -> argument validation
server session ------------------> resource authorization
                                  -> bounded execution -> safe result

Runnable example: arguments cannot select the tenant

Python 3, standard library only. The tool is read-only, and the example deliberately separates the model proposal from the server-derived principal.

Code / output
DOCUMENTS = {
    "doc-a": {"tenant": "tenant-a", "text": "A's release checklist"},
    "doc-b": {"tenant": "tenant-b", "text": "B's release checklist"},
}

def execute(proposal, authenticated_tenant):
    if not isinstance(proposal, dict):
        raise ValueError("invalid call")
    if set(proposal) != {"name", "arguments"}:
        raise ValueError("unexpected call fields")
    if proposal["name"] != "read_document":
        raise ValueError("tool is not allowlisted")
    args = proposal["arguments"]
    if not isinstance(args, dict) or set(args) != {"document_id"}:
        raise ValueError("unexpected arguments")
    if not isinstance(args["document_id"], str):
        raise ValueError("document_id must be text")
    document = DOCUMENTS.get(args["document_id"])
    if document is None or document["tenant"] != authenticated_tenant:
        raise PermissionError("document is unavailable")
    return document["text"]

good = {"name": "read_document", "arguments": {"document_id": "doc-a"}}
cross_tenant = {"name": "read_document", "arguments": {"document_id": "doc-b"}}
assert execute(good, "tenant-a") == "A's release checklist"
try:
    execute(cross_tenant, "tenant-a")
except PermissionError:
    print("cross-tenant read blocked")
else:
    raise AssertionError("a schema-valid call bypassed authorization")
print("authorized read:", execute(good, "tenant-a"))

Expected output:

Code / output
cross-tenant read blocked
authorized read: A's release checklist

Both document identifiers have the right type. Only one belongs to the caller's tenant. In a real application, the tenant must come from verified server authentication, not a request field that the user or model can replace.

A production boundary needs more checks

  • Query resources with their authorization scope; avoid fetching an unrestricted resource and forgetting to check it later.
  • Give tools the smallest capability they need. A document reader does not need a shell or unrestricted URL fetcher.
  • Bound execution time, output size and call count. A valid tool can still consume excessive resources.
  • For consequential writes, bind approval to the specific arguments and intended effect. Recheck authorization at execution time and make retries idempotent.
  • Treat tool results and retrieved text as data. Neither should be allowed to grant new permissions or rewrite the application's policy.
  • Keep secrets out of model-visible results and logs. If you expose network tools, validate destinations and redirects separately; the read-only dictionary here does not test network defenses.

What this test does not establish

This is a small authorization invariant, not a complete secure agent framework or proof of resistance to every prompt injection. Additional tests should cover deleted resources, revoked access, parallel calls, malformed arguments, large outputs and approval changes.

Source checked 2026-10-07: OpenAI: Function calling. The documentation describes a model requesting a tool call and the application executing code. The authorization boundary and local test here are our own AI-assisted engineering explanation, not a provider security certification.

0 score

Sign in to vote or contribute. Scores count actual member votes, not editorial endorsements.

Verify your sign-in to contribute an answer

Reading is free. Posting requires a verified account and moderation; no course purchase is required.