Structured outputs and tool calling: making LLMs reliable inside business workflows
Free text is hard to automate. JSON schemas, validation, tool calling and approval steps turn language models into dependable components of business systems, from email triage to booking changes.

Structured outputs and tool calling are the two techniques that turn a language model from a text generator into a component other software can rely on. A model that answers in prose is useful to a person reading it. A model that returns validated JSON, or asks the application to call a specific function with specific arguments, can sit inside an automated workflow — classifying requests, extracting data, and preparing actions for code or people to approve.
Why free text breaks automation
Ask a model to “find the booking reference in this email” and you might get the reference, a sentence containing it, two candidates, or an apology. Parsing that with regular expressions is fragile, and it fails silently. Business workflows need fields with known names, types and allowed values — and a clear way to say “not found”.
Structured outputs
The first step is to define the output as a schema. For a customer-email triage step, for example:
intent: one ofchange,cancel,refund_status,complaint,other,booking_reference: string or null,travel_date: date or null,urgency:low,normalorhigh,summary: a short string for the agent.
Many model APIs can now constrain responses to a supplied JSON Schema. Where that is not available, describe the schema in the instructions and parse strictly. Either way, the design choices matter:
- Use enums for anything the code will branch on. Never let the model invent a new category.
- Make “unknown” explicit with nullable fields, so the model is not pushed to guess.
- Keep schemas small. One focused extraction per call is more reliable than a giant object.
- Ask for evidence where useful — for example, the text span that contains the booking reference — so a reviewer can check quickly.
Validate everything
A schema-conforming response can still be wrong. Validation belongs in ordinary code:
- Parse and check the structure against the schema.
- Apply business checks: does the booking reference match a known format, and does it exist? Is the date in the future?
- If validation fails, retry once with the error message included, then route to a person rather than retrying indefinitely.
The model proposes; deterministic code decides whether the proposal is acceptable.
Tool calling
With tool calling, the application describes functions the model may use — their names, descriptions and argument schemas. The model responds with a request to call one, the application runs it, and the result goes back to the model. This is how an assistant looks up a booking before answering, rather than guessing.
Design tools like a careful API, because that is what they are:
- Narrow and specific —
get_booking(reference)is better thanquery_database(sql). - Clear descriptions — the model chooses tools from their descriptions, so write them precisely, including when not to use the tool.
- Safe defaults — limit result sizes, scope access to the current customer or account, and return errors the model can understand.
- Idempotent writes — any tool that changes state should accept an idempotency key, for the same reasons described in idempotency in booking and payment APIs.
Separate reading from acting
The most important design decision is which tools need approval.
- Read tools — look up a booking, check a refund status, search a policy document — can usually run automatically.
- Write tools — change a booking, issue a refund, send a customer message — should produce a proposed action that code validates against business rules and, where the impact is significant, a person approves.
A practical pattern is a “draft” tool: the model calls propose_refund(order, amount, reason), which creates a pending item in an operations queue with all the context, instead of moving money. This keeps people in control of the decisions that matter, the same principle behind AI agents in travel and back-office automation.
Guard against untrusted input
Inputs such as customer emails, supplier notices and documents are untrusted. They may contain text that looks like instructions to the model. Defences are mostly architectural:
- tools are scoped to what the current task needs, nothing more,
- the model cannot grant itself new permissions,
- actions are validated by code against the authenticated user and account, not against what the text claims,
- high-impact actions always pass through approval.
Treat the model’s output with the same suspicion you would apply to user input.
Measure it like any other component
Structured output makes evaluation straightforward: fields can be compared with expected values. Build a test set from real, anonymised examples, track accuracy per field and per intent, and run it on every prompt or model change. How to evaluate LLM applications covers building that test set. In production, log inputs, outputs, validation failures and human corrections — the corrections are the best source of new test cases.
Summary
Define outputs as small schemas with enums and explicit unknowns, validate them with ordinary code, expose narrow and well-described tools, let read tools run while write tools propose actions for validation and approval, treat all input text as untrusted, and evaluate continuously. Used this way, a language model becomes a dependable step in a business workflow rather than an unpredictable one.
Frequently asked questions
What are structured outputs in LLMs?
Structured output means asking a language model to return data in a defined format, usually JSON that matches a schema, instead of free text. Many model APIs can constrain responses to a supplied schema, which makes outputs easier to validate and use in code.
What is tool calling?
Tool calling, also called function calling, lets a model request that the application run a defined function with specific arguments, such as looking up a booking. The application executes the function and returns the result to the model, which then continues.
Should an LLM be allowed to take actions directly?
Read-only tools can usually run automatically. Actions that change money, bookings or customer data should be validated by code against business rules and, where the stakes are high, approved by a person before they execute.