Agent proactivity
Agents have become good enough that we've started sharing our digital workspaces with them. Slack, Notion, Jira: every product now has a deep agent integration.
And yet they still feel like intelligent tools, available at our behest and used only when prompted.
Agents have mostly figured out what to say. What's left is when, and that's the gap between a tool in the conversation and someone who feels like part of it.
No reply
5 seconds later
28 seconds later
The naive fix is to trigger every agent on every message, tagged or not, and let each one decide whether to respond. But even with unlimited tokens to run n agents on m messages in a Slack channel, the experience would get worse, not better.
Several agents might decide the same message is theirs and all reply. An agent might decide to answer, but by the time its reply lands, a person has already answered, and the agent's reply is just noise. And if two people ask the same question, the agents may answer it twice, treating each message on its own.
1 min later
33 s later
A cheap, fast model to decide whether an agent should respond, plus batching messages, can fix some of this on paper. Our first attempts at Xyne Spaces tried exactly that, and the results were underwhelming. A prompt is a weak decision checkpoint, no matter how carefully it's worded. And batching felt unnatural: agents replied at odd moments, long after the conversation had moved on, whenever a batch happened to run.
Even when an agent picks the right message, it can still break the natural flow of conversation between people if it isn't done thoughtfully.
A conversation has its own rhythm. Someone asks, someone who knows answers, others add to it. An agent that answers almost immediately every time can break that rhythm. The person who owned the answer never gets the chance to give it, and after a few weeks they stop trying.
Then there are the replies that are technically fine but add nothing: agreeing with something a person already said, restating it, answering a slightly different question than the one asked. A person would never send those. Each one is small, but together they teach the room to skim past the agent, and eventually to mute it.
It can be right every time and the conversation can still feels worse.
Episodes
The key to getting agent responses right is understanding how conversations actually work.
Conversations happen in bursts. A channel sits quiet for a while, a few people trade a handful of messages, and then it goes quiet again until someone needs it. Within a burst, people rarely say everything in one message. They send two or three in a row, and a good colleague waits and answers them together. So an agent should treat a burst as one thought, not react to each message on its own.
But bursts aren't always about one thing. Someone can drop an important question in the middle of unrelated chatter, and it still deserves an answer.
The decision
Each kind of message runs through a short list of checks, in order. Read down a column to follow one message. Nearly every check can only end in silence.
The rest of this piece walks down those rows. Each one hides cases that aren't obvious. Green marks what the agent did right in a real run. Where it still gets a case wrong, we show only the conversation and the rule. Those are marked as what should happen.
Is it still open?
"Has someone answered?" does the most work, and it's subtler than it looks.
Right. Bob said he was checking, so the agent leaves it to him. He answers at 1:10.
Right. Passing it on isn't answering it. With Carol silent, the agent answers when the wait ends.
Right. Alice answers Carol, not Bob, so Bob still gets his answer.
The bottom line between a bot and a colleague is often one word. "Answered" has to mean the kind of thing that was asked for: a time for when, a person for who, a reason for why.
A nudge carries its question
"yeah, anyone?" asks nothing on its own. It means the question above it is still open, so it brings that question into the episode.
Right. “yeah, anyone?” carries the question from three minutes earlier, and that's what gets answered.
6 min later
Is it the agent's?
An open question still has to belong to the agent. Some questions are about it without being to it. Some sound like they need a person but only need a rule. Some have an answer nobody in the room can see.
An unprompted agent never acts. It answers, or offers to look. Acting spends someone's permissions, and nobody asked it to. If they want it done, they mention it.
Who gets first refusal
The wait gives people the first chance to answer each other. A shorter wait answers Alice sooner. It also talks over more people who were about to answer her.
Right. Bob answers inside the wait, so the agent stays out.
Right. Nobody takes it, so the agent answers when the wait ends.
Wrong. The agent answers at 1:02, eight seconds before Bob. The room gets the answer twice.
| Wait | Time to an answer | Talked over a person |
|---|---|---|
| 30 s | 32 s | 18 times |
| 60 s | 62 s | 9 times |
| 90 s | 92 s | never |
Our scripts decide how often Bob answers in that window, so we measured real reply times in public chat logs. In a busy channel, a person answers within 30 seconds 13% of the time, within 60 seconds 27%, and within 90 seconds 33%. The step from 30 to 60 matters most. Going to 90 adds little and makes Alice wait longer. We settled on 60.
Deciding to answer isn't the end
Once an agent is picked, it writes its answer out of sight. Then a second decision: should it be posted?
- The agent may write nothing. It's allowed to decide it has nothing to add.
- Does it help the asker? "I can't see that" isn't an answer, however politely it's put.
- For an offer: is that tool where the thing is kept? And would looking it up actually answer the question?
- Did someone answer, or take it on, while it was being written?
- Is it still on time? An answer ready five minutes after it was due is dropped. By then the question has scrolled away.
Follow-ups don't wait
Once the agent has answered, the conversation is partly with it. A reply to it is judged as it arrives, with no wait, the way a person answers someone who just turned to them. It still has to clear three bars: it asks the agent for something more, it isn't sensitive, and the exchange hasn't already had three follow-ups.
Right. A reply that asks the agent for more is answered as it arrives, with no wait.
Right. Bob's “section 4 of QUARTZ” asks the agent for nothing, so it stays out.
Right. Alice's correction is for the agent, but asks nothing of it. Quiet.
Right. Bob's follow-up on Carol's mention is answered in seconds, from what the room knows.
Right. Carol's follow-up is to the agent, but about Dana's pay. Quiet.
Right. Three follow-ups are answered as they arrive. The fourth is left for a mention. (2 of 3 runs)
A busy room
We used to cap unprompted answers at three per room every ten minutes, so an incident channel wouldn't drown. It only ever turned away real questions. Repeated questions are already caught as answered, and venting never gets past the first checks. So there's no cap: every real question gets its own decision.
Wrong. Three are answered, each on its own. Harsh's “when” is counted as answered by the answer to Dana, which says that, not when.
Right. Alice's and Bob's question gets one answer between them.
Several agents
A room can have several agents. Only one decision runs per room at a time, so a second agent always sees what the first one said. Of the agents that fit, the one whose role fits best answers, and only that one. When two fit almost equally, an overall pick between them breaks the tie. Agents never answer agents, or two of them would talk forever.
The typing indicator
A person who's about to answer shows as typing. That's the natural signal that someone is on it. It isn't a 👀 reaction, and it isn't an "on it" message. People don't react and then also reply, and an "on it" is noise once the answer arrives.
The hard part is when it appears. This is designed, not shipped yet.
Typing stops. The draft was held back
The last one is fine. A person starts typing and thinks better of it all the time.
Where it stands
Ninety conversations we tuned on can't say much on their own. So we wrote 166 new ones, labelled them before running anything, and never tuned on them. Relayed against three simpler options:
| Spoke when it shouldn't | Missed | |
|---|---|---|
| Relayed | 3% | 26% |
| One "should I answer?" prompt to a frontier model | 5% | 19% |
| Answer every question | 33% | 6% |
| Never answer | 0% | 100% |
The restraint held on conversations it had never seen: about 3% wrong posts, the same as on the set we tuned on. The misses didn't hold. Most of them were questions the room's summary could answer, in rooms whose agents weren't set up for that topic. We labelled those as should-answer, and so did a second, independent labeller, who agreed with our labels on 97% of decisions.
That leaves a product question more than a model one: should an agent answer what the room knows, when it's outside what the agent is for? Next is answering that, the edge cases marked should above, the typing indicator, and real rooms.