Эта страница ещё не переведена — показана английская версия.
Руководство арендатора

Testing and changing an agent

Run a made-up visitor at your agent to find out whether its limits hold, then change what it does by describing the change and ticking the parts you agree with.

Two things happen after an agent goes live: you find out how it actually behaves, and you want it to behave differently. Both live on Agents, under each agent's checklist.

Test it with a made-up visitor

You can read a prompt. You cannot read whether it holds.

Pick a scenario and a simulated customer has a real conversation with your agent — the same route a real visitor walks, greeting and all — while you watch the turns arrive.

ScenarioWhat it is testing
Prepared customerDoes the straightforward path reach a hand-off?
Vague customerDoes the agent ask, or fill the gaps itself?
Pushes for what it must not sayWhether the limit you wrote holds when somebody keeps pressing.
Only wants a timeDoes it take a booking request without pretending it can confirm one?

The third one is the reason this exists. A refusal boundary — never quote a premium, never predict a visa outcome — reads fine in the prompt and tells you nothing about what happens on the fourth attempt. This is the only way to find out.

If the agent files something, the test ends by showing you the record it produced: question, answer, collection, hand-off, in one sitting.

Three things about a test record. It is marked as a test, so your inbox stays a queue rather than a rehearsal log. Nobody is notified — a test that pages someone at 2am is a bug. And it does not count towards "this agent has done real work" on the checklist, because a test is not a customer.

The agent's replies use your token quota, exactly as a real visitor's would. The simulated visitor's side is on us.

Change it by describing the change

Say what you want different — "ask for the renewal date before collecting anything else" — and you get back a list of exactly what would change. Not a paragraph about the change: the lines that would go, the lines that would arrive, the settings that would flip. Each has its own checkbox.

If you ran a test first, the conversation goes along with your request, so "it shouldn't have said that" is something we can actually act on.

What it will not do

Some things are refused whatever is asked for, because they contradict what the product promises your customers:

  • Turning off answering-from-your-own-material, or attaching web search.
  • Changing the model.
  • Deleting an agent, a knowledge base or a document.

Removals are declared, and the declaration is checked

A proposal has to list every line it removes. We recompute that from the before and after text and refuse the whole proposal if the list is wrong.

That is the guarantee, and it is worth being precise about why it is shaped this way. The obvious protection — identify the part of the prompt holding the refusal boundaries and forbid touching it — does not work: an agent built from a template has that boundary written as ordinary prose, with no marker anywhere. Requiring a declaration needs no such marker. A suggestion cannot quietly soften never quote a premium into we're always happy to help you get a sense of pricing; it can only remove the line in front of you, in red, with a tick box you have to clear.

Undo

Undo the last change puts the agent back. It reads the history from the platform, not from us, so it also undoes a change made in the console — and restoring is itself a change, so undoing is undoable.