Proposal: report RTP inactivity to ARI/AMI instead of hanging up (rtp_timeout)

Hello,

Following a suggestion on GitHub issue #2202
([new-feature]: res_pjsip_sdp_rtp: optionally report RTP inactivity to the channel's application instead of hanging up · Issue #2202 · asterisk/asterisk · GitHub), I would like input from
people who might use this — especially those building real-time voice agents
on ARI.

The problem. An endpoint’s rtp_timeout can only hang up: after N seconds with
neither RTP nor RTCP, rtp_check_timeout() soft-hangs-up the channel. An ARI
application has no other way to learn that a channel’s media has stopped:
there is no event for it, and rtp_statistics exposes RTP only, so it cannot
tell a far end that suppresses silence from a dead path.

Why it matters for voice agents. A voice agent decides when to speak from the
caller’s silence. Many trunks send no RTP at all while the caller is silent.
Today the application cannot tell “the caller is thinking” from “the network
is gone”: the agent either waits forever or keeps asking “are you there?” into
a dead line — and the call is billed until something hangs it up. With a
mobile caller, a short network drop should pause the agent, not end the call
and lose the conversation; when media returns, the agent should carry on.

The proposal. Keep hang-up as the default and add an option, e.g.
rtp_timeout_action = hangup | notify. With notify, Asterisk raises a
channel event when the timeout expires (e.g. ChannelMediaInactive) and
another when RTP or RTCP arrives again (e.g. ChannelMediaActive), and leaves
the decision to the application — pause, hold for a while, tell the other
participants, or hang up. No ABI change; existing configurations behave as
now. It would be built on top of #2143, which moves the check to the session
level.

Our side: we are building a framework that drives Asterisk through ARI for AI
voice agents (receptionist / secretary), and we are ready to implement this
with a testsuite test once the shape is agreed.

Questions where your view would help:

  1. Would you use this, and for what?
  2. Endpoint option, or would you rather set it per channel?
  3. Event names, and should AMI get the events too?

This message was drafted with the help of an AI assistant; I have reviewed it.

This caught my eye as this statement is completely counter to what I’ve seen throughout all my years. This is referred to as silence suppression and is not something really supported by Asterisk, and isn’t something I see these days from providers or even endpoints.

Thank you — you’re right to call that out. That sentence was our assumption,
not something we have observed from a provider; I withdraw it.

That leaves the case we care about most: a caller whose network disappears
without a BYE (mobile coverage lost, NAT binding dropped). RTP then stops
entirely. If RTP normally keeps flowing through silence, an application can
already detect this by polling rtp_statistics for each channel — is that the
approach you would recommend?

Our hesitation is that it means a periodic request per participant, while
Asterisk already tracks exactly this for rtp_timeout and can only act on it
by hanging up. An event would let the application decide — pause the agent,
hold the participant for a short window, then hang up — without polling.

If you think polling is the right answer, we’ll go that way.

Silence and non-flowing RTP are two separate things.

Silence would not appear in RTP statistics because the RTP stream itself is still flowing albeit with silence within it. Either the voice agent should handle this, or TALK_DETECT.

The only time RTP timeout applies is generally with actual directly connected endpoints and having those drop out. It can certainly apply if an upstream provider has some fundamental issue and goes down, or your internet goes down, but that’s more rare. I don’t believe I’ve ever seen it trigger when you are connected to an upstream provider with a mobile call on the other side and they encounter issues.

Have you actually created the scenario(s) and looked at what actually happens with upstreams and such?

I suppose what I’m getting at is:

Would this actually solve your problem, or is this idea predicated on assumptions and the work wouldn’t actually solve it?

Honestly: for the upstream case it was predicated on assumptions, and your
replies show those assumptions were wrong. We had only tested two Asterisk
instances connected by a SIP trunk on one machine, with one of them frozen —
not a real provider with a mobile call on the other side.

The case where we really need this is directly connected endpoints: WebRTC
legs from browsers (and softphones) whose network drops for a few seconds.
Even there we don’t know yet that an event would be enough — bringing the
browser back into the same channel would also need an ICE restart, which we
haven’t tested.

So no, we can’t claim it would solve our problem. We’ll measure the browser
case, and our upstream once it is in place, and come back only with results.
Until then please consider the proposal on hold — I’ll note that on #2202.
Thank you for pushing on this; it saved us from building the wrong thing.