4 minute read · Practical field guide
Define the action before judging the answer
Choose a narrow operation for the first workflow, such as proposing a new appointment time without confirming it. List the fields needed to complete that operation: customer identity, appointment identifier, requested date, relevant time zone and any constraints. Specify which fields the assistant may infer from reliable context and which require an explicit check. Avoid defining success as merely producing a fluent reply.
Use a small state model: request received, context incomplete, awaiting clarification, ready for validation, action attempted and outcome recorded. A human handoff is another deliberate outcome. These states are illustrative application design, not provider delivery statuses. Keeping them separate prevents a message marked delivered from being mistaken for an appointment successfully changed.
Walk through an ambiguous request
Consider a fictional customer with a site inspection on Tuesday and an equipment delivery on Wednesday. They reply, “Can we move that to Friday?” A model may assign high confidence to the idea of rescheduling, but the target remains ambiguous. Ask: “Do you mean Tuesday’s inspection or Wednesday’s delivery?” Store the clarification against the open request rather than creating a new unrelated task.
If the customer answers “the delivery,” you have resolved one field, not necessarily all of them. You may still need a time window and a capacity check. The assistant should ask the next necessary question or offer verified choices. Do not invent availability to keep the conversation moving. When the business cannot establish a safe target, route the request to a person with the original text and the candidates clearly listed.
Use confirmation to summarize a specific proposal
A confirmation is useful when it names exactly what would change: “I can request moving delivery 42 from Wednesday morning to Friday afternoon. Is that the change you want?” Its wording should match the assistant’s actual authority. If the assistant can only request a change, it must not announce that the schedule has already been updated.
Confirmation does not replace identity checks or business rules. A person replying from a known number may still be asking for an action that requires another verification step. Define these requirements in application policy rather than improvising them from chat. Recheck current record state before execution so a colleague’s intervening change does not get overwritten by an older conversation.
Carry context across replies without guessing
Attach each incoming answer to an identifiable pending question. Keep the original request, the last question asked, any established values and a reasonable expiry condition. If a customer returns days later with “yes,” check whether the proposal is still valid. A stale confirmation should lead to a fresh check rather than silently executing yesterday’s plan.
Allow the user to change topic or correct an earlier value. “Actually, keep Wednesday” should update or cancel the pending request, not be treated as another answer in a fixed sequence. Duplicate event delivery should leave the workflow unchanged after the first successful processing. Use durable identifiers and explicit state transitions so retries cannot create additional bookings or repeated outgoing confirmations.
Make escalation a complete handoff
When the assistant cannot continue, hand over enough information for a colleague to act: the customer’s exact request, relevant record links, clarification already attempted, fields still missing and any execution result. State whether a change was applied, merely proposed or not attempted. A human should not have to guess whether the assistant already promised something to the customer.
The customer-facing response should also be accurate. Explain that the team will review the request and give only a response expectation the business can meet. Do not keep sending clarification questions after a human has taken ownership. Establish how automation pauses, how staff record a resolution and how the assistant is allowed to resume. A handoff without ownership is simply a new unattended queue.
Test the cases that should slow the assistant down
Build a small evaluation set from fictional or appropriately handled examples: two matching orders, a relative date across time zones, a correction after confirmation, a duplicate webhook, a reply to an expired proposal and a failed business-system update. Write the expected action before running each case. Some cases should ask a question or stop for a human; do not score those as failures merely because no automatic change occurred.
Inspect the full record after each test: incoming event, interpreted request, validation decision, attempted operation and outgoing explanation. Measure wrong actions separately from unresolved conversations and transport failures. Public messaging documentation explains the channel and event mechanics; it does not prove the quality of your assistant’s decisions. The diagram above is a starting design to validate in your own application, with limited scope and explicit ownership.