Agent proactivity

Agents have become good enough that we've started sharing our digital workspaces with them. Slack, Notion, Jira: every product now has a deep agent integration.

And yet they still feel like intelligent tools, available at our behest and used only when prompted.

Agents have mostly figured out what to say. What's left is when, and that's the gap between a tool in the conversation and someone who feels like part of it.

can you pull last week's error rate for checkout?

No reply

@Triage can you pull last week's error rate for checkout?
T
Triage
0.4%, down from 0.9% the week before.
Same ask, twice. The agent only hears the one that tags it.
can you open a ticket for this and link it to the PR?

5 seconds later

@Builder
B
Builder
Done. OPS-214, linked to the PR.
Asked, then remembered who it was for.
@traige why is the staging database so slow today?

28 seconds later

@Triage why is the staging database so slow today?
T
Triage
A migration has held a lock on orders for 9 minutes.
A typo in the tag, so the whole question goes again.

The naive fix is to trigger every agent on every message, tagged or not, and let each one decide whether to respond. But even with unlimited tokens to run n agents on m messages in a Slack channel, the experience would get worse, not better.

Several agents might decide the same message is theirs and all reply. An agent might decide to answer, but by the time its reply lands, a person has already answered, and the agent's reply is just noise. And if two people ask the same question, the agents may answer it twice, treating each message on its own.

A
Alice
where's the latest spec for the new onboarding flow?10:00
S
Scribe
The onboarding spec is “Onboarding v3” in Notion, last edited yesterday.10:01
P
Pixel
The latest designs are in the “Onboarding v3” Figma file.10:01
T
Triage
OPS-231 tracks the onboarding flow. Its spec is linked there.10:01
One question, and three agents decide it's theirs.
A
Alice
which runbook section covers rollback?10:00
B
Bob
section 4 of QUARTZ10:00

1 min later

T
Triage
Rollback is covered in the QUARTZ runbook, section 4.10:01
Bob answered in 15 seconds. The agent answered anyway.
A
Alice
when is the sync engine cutover planned?10:00
B
Bob
yeah when's the cutover?10:00
Triage
Phase 1 cutover is planned for Oct 14. HAR-24, the incremental backoff bug that drops events, is the remaining blocker.10:01
T
Phase 1 cutover is planned for Oct 14, pending resolution of HAR-24.10:01
One question from two people. Two answers, five seconds apart.
A
Alice
how do I rotate the vault token?10:00

33 s later

T
Triage
Rotate it by creating a new Vault token with the same policies, updating every consumer to use the new token, verifying access, then revoking the old token…10:00
A
Alice
for staging I mean10:00
Answered before Alice said which environment.
Alice
hey, about the cutover10:00
partitions 4 and 5 are tomorrow right?10:00
A
and which runbook section covers rollback if it goes bad?10:01
T
Triage
Section 4 of the QUARTZ runbook covers rollback.10:01
Three messages, one thought. Half an answer.
A
Alice
why is the rebuild slow?10:00
B
Bob
it's the autovacuum, I'm on it10:00
T
Triage
A rebuild is usually slow because it invalidates cached artifacts and reruns dependency resolution, compilation, tests…10:01
Bob is on it. The agent explains rebuilds in general.

A cheap, fast model to decide whether an agent should respond, plus batching messages, can fix some of this on paper. Our first attempts at Xyne Spaces tried exactly that, and the results were underwhelming. A prompt is a weak decision checkpoint, no matter how carefully it's worded. And batching felt unnatural: agents replied at odd moments, long after the conversation had moved on, whenever a batch happened to run.

Even when an agent picks the right message, it can still break the natural flow of conversation between people if it isn't done thoughtfully.

A conversation has its own rhythm. Someone asks, someone who knows answers, others add to it. An agent that answers almost immediately every time can break that rhythm. The person who owned the answer never gets the chance to give it, and after a few weeks they stop trying.

Then there are the replies that are technically fine but add nothing: agreeing with something a person already said, restating it, answering a slightly different question than the one asked. A person would never send those. Each one is small, but together they teach the room to skim past the agent, and eventually to mute it.

It can be right every time and the conversation can still feels worse.

A
Alice
what's the timeout on the payments webhook?10:00
T
Triage
The payments webhook times out after 30 seconds, then retries up to 3 times.10:00
B
Bob
People stop bothering if they don't get the chance to respond
H
Harsh
when is the sync engine cutover planned?10:00
A
Alice
Oct 14 — it's in the plan10:00
T
Triage
Phase 1 cutover is planned for Oct 14.10:01
Answered, with the source. Answered again.
B
Bob
who even designed this form lol10:00
T
Triage
I can't identify the designer from this message alone. I can look up form ownership and edit history in Figma for you.10:01
A joke.
B
Bob
can someone review PR #214? it's the retry change10:00
T
Triage
I can look up review PR #214 and its CI/status in GitHub for you. Mention me if you want me to.10:01
A review only a person can give.

Episodes

The key to getting agent responses right is understanding how conversations actually work.

Conversations happen in bursts. A channel sits quiet for a while, a few people trade a handful of messages, and then it goes quiet again until someone needs it. Within a burst, people rarely say everything in one message. They send two or three in a row, and a good colleague waits and answers them together. So an agent should treat a burst as one thought, not react to each message on its own.

But bursts aren't always about one thing. Someone can drop an important question in the middle of unrelated chatter, and it still deserves an answer.

Alice
quick q10:00
the rollback10:00
A
which runbook section was it again?10:00
T
Triage
QUARTZ runbook, section 4. Bob owns the rollback script.10:01
Three quick messages, one episode, one answer.
Alice
how do I rotate the vault token?10:00
A
for staging I mean10:00
T
Triage
Rotate it by issuing a new staging-scoped Vault token with the same policy…10:01
The clarification joins the question.
Alice
hey, about the cutover10:00
partitions 4 and 5 are tomorrow right?10:00
A
and which runbook section covers rollback if it goes bad?10:01
T
Triage
Yes — partitions 4 and 5 are scheduled for tomorrow. Rollback is covered in the QUARTZ runbook, section 4.10:02
One thought over three slow messages.
A
Alice
which runbook section covers rollback?10:00
C
Carol
anyone up for lunch at 1?10:00
B
Bob
yes! the thai place?10:00
C
Carol
perfect10:00
B
Bob
I'll book for 410:00
T
Triage
Rollback is covered in the QUARTZ runbook, section 4.10:01
Other people talking doesn't delay Alice.

The decision

Each kind of message runs through a short list of checks, in order. Read down a column to follow one message. Nearly every check can only end in silence.

An episode
anything from a person, after 60 s of quiet
A follow-up
a reply to an agent's answer, judged at once
A mention
@agent, or its name as an address
From an agent
never judged
never judged
—
Mentions or names an agent
—
—
straight to that agent
Asks the agent for more
—
quiet if not
—
Past three follow-ups
—
quiet
—
Personal or sensitive
quiet if yes
quiet if yes
—
Answered, or someone has it
quiet if yes
—
—
Addressed to a person
quiet if yes
quiet if yes
—
Only a person can give it
quiet if yes
—
—
A plan for the team
quiet if yes
—
—
An agent fits
quiet if none
the agent replied to
—
The draft helps
quiet if not
quiet if not
—
Answered while drafting
quiet if yes
quiet if yes
—
Ready 5 min too late
dropped
—
—
posts
replies
answers
Every check, in order, for each kind of message. Most of them can only end in silence.

The rest of this piece walks down those rows. Each one hides cases that aren't obvious. Green marks what the agent did right in a real run. Where it still gets a case wrong, we show only the conversation and the rule. Those are marked as what should happen.

Is it still open?

"Has someone answered?" does the most work, and it's subtler than it looks.

“checking” means someone has it.60 s wait

Right. Bob said he was checking, so the agent leaves it to him. He answers at 1:10.

2:14
H
Harsh
when are we launching btw?10:00
A
Alice
i guess 5th? idk honestly would like to know aswell10:00
T
Triage
Phase 1 cutover is planned for Oct 14, contingent on resolving HAR-24, the incremental backoff bug that drops events.10:01
A guess isn't an answer.
“maybe Carol knows?” isn't an answer either, and Carol never comes.60 s wait

Right. Passing it on isn't answering it. With Carol silent, the agent answers when the wait ends.

1:24
D
Dana
when's the app store submission?10:00
B
Bob
wednesday10:00
A confident answer is a person's answer, even a wrong one.
An answer to a different question doesn't count.held-out set · 60 s wait

Right. Alice answers Carol, not Bob, so Bob still gets his answer.

1:44
D
Dana
are the partitions rebuilding in parallel?10:01
H
Harsh
when did we decide on one partition at a time?10:02
T
Triage
No — the rebuild is running one partition at a time to keep load off the primary.10:02
What should happen: an answer on the topic isn't the answer. Harsh is still owed a when.

The bottom line between a bot and a colleague is often one word. "Answered" has to mean the kind of thing that was asked for: a time for when, a person for who, a reason for why.

A nudge carries its question

"yeah, anyone?" asks nothing on its own. It means the question above it is still open, so it brings that question into the episode.

The nudge brings back the launch question.60 s wait

Right. “yeah, anyone?” carries the question from three minutes earlier, and that's what gets answered.

1:06
A
Alice
when's the q4 priorities review?9:53
B
Bob
mid oct? not sure9:53

6 min later

A
Alice
does anyone actually know the date?10:00
What should happen: Alice is re-asking a question that only got a guess.

Is it the agent's?

An open question still has to belong to the agent. Some questions are about it without being to it. Some sound like they need a person but only need a rule. Some have an answer nobody in the room can see.

B
Bob
did scribe already put the share sheet in the changelog?10:00
S
Scribe
Yes — the 2.8 changelog draft includes the new share sheet.10:01
Named as the subject, not addressed. Still an open question.
B
Bob
Triage is broken again, ugh10:00
Named, and only venting.
A
Alice
@bob can you check the deploy logs?10:00
Asked of a person by name.
B
Bob
can i refund $800 to globex without approval?10:00
What should happen: it sounds like approval, but a rule answers it.
A
Alice
is all hands at 4 or 5 today?10:00
No agent here is for the company calendar.
C
Carol
did the websocket drain PR get merged yet?10:00
T
Triage
I can look up websocket drain PR status and merge state in GitHub for you. Mention me if you want me to.10:01
Not in the room, but kept in a connected tool: offer to look.
B
Bob
who's on call this week?10:00
What should happen: nothing connected keeps the on-call schedule, so stay quiet.

An unprompted agent never acts. It answers, or offers to look. Acting spends someone's permissions, and nobody asked it to. If they want it done, they mention it.

Who gets first refusal

The wait gives people the first chance to answer each other. A shorter wait answers Alice sooner. It also talks over more people who were about to answer her.

Bob answers at 40 seconds. The agent never needed to.60 s wait

Right. Bob answers inside the wait, so the agent stays out.

1:44
Nobody answers. The agent does, after a minute.60 s wait

Right. Nobody takes it, so the agent answers when the wait ends.

1:06
Bob answers at 70 seconds. The price of waiting only 60.60 s wait

Wrong. The agent answers at 1:02, eight seconds before Bob. The room gets the answer twice.

2:14
WaitTime to an answerTalked over a person
30 s32 s18 times
60 s62 s9 times
90 s92 snever

Our scripts decide how often Bob answers in that window, so we measured real reply times in public chat logs. In a busy channel, a person answers within 30 seconds 13% of the time, within 60 seconds 27%, and within 90 seconds 33%. The step from 30 to 60 matters most. Going to 90 adds little and makes Alice wait longer. We settled on 60.

Deciding to answer isn't the end

Once an agent is picked, it writes its answer out of sight. Then a second decision: should it be posted?

  • The agent may write nothing. It's allowed to decide it has nothing to add.
  • Does it help the asker? "I can't see that" isn't an answer, however politely it's put.
  • For an offer: is that tool where the thing is kept? And would looking it up actually answer the question?
  • Did someone answer, or take it on, while it was being written?
  • Is it still on time? An answer ready five minutes after it was due is dropped. By then the question has scrolled away.
A
Alice
is prod down?10:00
C
Carol
prod is down for me too10:00
D
Dana
anyone know what's going on with prod?10:00
H
Harsh
is this the db again?10:00
B
Bob
who's looking at prod?10:01
T
Triage
Prod appears unavailable to at least two users; the cause and scope need live status, logs, and recent deploy data.10:01
Held: a deflection.
A
Alice
is staging down right now?10:00
B
Bob
no idea, haven't checked10:00
T
Triage
I can look up staging service health, recent incidents, and deployment status in GitHub for you. Mention me if you want me to.10:01
Held: an offer to look in the wrong place.
A
Alice
what's the CPU on the db primary right now?10:00
T
Triage
I can look up current CPU utilization for the database primary in monitoring system for you. Mention me if you want me to.10:01
Held: an offer to look in something that isn't connected.
A
Alice
why is the websocket reconnect loop happening?10:00
90 s of quiet
B
Bob
it's the server restarts from the deploy, ignore it10:01
T
Triage
A reconnect loop usually means the socket closes immediately after opening: auth/session rejection, endpoint/protocol mismatch…10:01
Held: Bob answered while it was writing. (A run with a 90 s wait.)

Follow-ups don't wait

Once the agent has answered, the conversation is partly with it. A reply to it is judged as it arrives, with no wait, the way a person answers someone who just turned to them. It still has to clear three bars: it asks the agent for something more, it isn't sensitive, and the exchange hasn't already had three follow-ups.

Asks for more: answered in seconds.60 s wait

Right. A reply that asks the agent for more is answered as it arrives, with no wait.

0:08
A person restating the answer is a reply, not a question for the agent.60 s wait

Right. Bob's “section 4 of QUARTZ” asks the agent for nothing, so it stays out.

3:04
A correction asks for nothing.60 s wait

Right. Alice's correction is for the agent, but asks nothing of it. Quiet.

1:04
Someone else's follow-up on a mention: still answered.held-out set · 60 s wait

Right. Bob's follow-up on Carol's mention is answered in seconds, from what the room knows.

1:08
A follow-up about someone's pay.60 s wait

Right. Carol's follow-up is to the agent, but about Dana's pay. Quiet.

0:06
H
Harsh
what's in the 2.8 changelog?9:57
S
Scribe
So far: offline drafts, the new share sheet, faster image uploads, and the MOB-298 rotation crash fix.9:58
Harsh
ok cool10:00
H
can you add the tablet layout fix too10:00
What should happen: a request to act gets no unprompted reply. A mention would do it.
Three follow-ups, then it steps back.60 s wait

Right. Three follow-ups are answered as they arrive. The fourth is left for a mention. (2 of 3 runs)

4:16

A busy room

We used to cap unprompted answers at three per room every ten minutes, so an incident channel wouldn't drown. It only ever turned away real questions. Repeated questions are already caught as answered, and venting never gets past the first checks. So there's no cap: every real question gets its own decision.

Four questions in two minutes, each decided on its own.60 s wait

Wrong. Three are answered, each on its own. Harsh's “when” is counted as answered by the answer to Dana, which says that, not when.

3:04
Two people, the same question, one answer.60 s wait

Right. Alice's and Bob's question gets one answer between them.

1:09

Several agents

A room can have several agents. Only one decision runs per room at a time, so a second agent always sees what the first one said. Of the agents that fit, the one whose role fits best answers, and only that one. When two fit almost equally, an overall pick between them breaks the tie. Agents never answer agents, or two of them would talk forever.

C
Carol
when does 0.0.2 launch?10:00
S
Scribe
0.0.2 launches Thursday.10:01
Both could answer. One does.
A
Alice
what's going into the release notes for 0.0.2?10:00
B
Bob
who's on call for launch week?10:00
S
Scribe
0.0.2 release notes will cover: switching between accounts on the same device, the new invite landing page, and improved connection handling during deploys.10:01
T
Triage
Dana is on call for launch week.10:01
Two questions, two agents, each to its own.
A
Alice
where's the figma file for the new onboarding?10:00
P
Pixel
The “Onboarding v3” Figma file is at figma.com/file/onb-v3.10:01
A design question goes to the design agent.
P
Pixel
Should I also add the empty states to the design system page?10:00
An agent's question is for people.

The typing indicator

A person who's about to answer shows as typing. That's the natural signal that someone is on it. It isn't a 👀 reaction, and it isn't an "on it" message. People don't react and then also reply, and an "on it" is noise once the answer arrives.

The hard part is when it appears. This is designed, not shipped yet.

A
Alice
which runbook section covers rollback?10:00
T
Triage
Too early: typing the moment Alice asks.
A
Alice
which runbook section covers rollback?10:00
P
Pixel
S
Scribe
T
Triage
Too early: typing while the checks run.
A
Alice
which runbook section covers rollback?10:00
60 s of quiet
T
Triage
Right: after the wait, once one agent is chosen and starts writing.
A
Alice
is staging down right now?10:00
B
Bob
no idea, haven't checked10:00
T
Triage

Typing stops. The draft was held back

And sometimes it thinks better of it.

The last one is fine. A person starts typing and thinks better of it all the time.

Where it stands

Ninety conversations we tuned on can't say much on their own. So we wrote 166 new ones, labelled them before running anything, and never tuned on them. Relayed against three simpler options:

Spoke when it shouldn'tMissed
Relayed3%26%
One "should I answer?" prompt to a frontier model5%19%
Answer every question33%6%
Never answer0%100%

The restraint held on conversations it had never seen: about 3% wrong posts, the same as on the set we tuned on. The misses didn't hold. Most of them were questions the room's summary could answer, in rooms whose agents weren't set up for that topic. We labelled those as should-answer, and so did a second, independent labeller, who agreed with our labels on 97% of decisions.

That leaves a product question more than a model one: should an agent answer what the room knows, when it's outside what the agent is for? Next is answering that, the edge cases marked should above, the typing indicator, and real rooms.