Manager Performance Improvement PR

I’ve been poking at manager performance for the last two months or so experimenting with things and have put up a pull request:

Which changes how event queueing in manager works.

For developers, or even users, who are curious how it worked it used a single list of events. Each session would know its current position in the list, wake up, move through the list (either sending or not sending), stop, rinse and repeat. Another thread would wake up every so often and clean up any events that were no longer needed.

This new change makes it so that each session has its own queue. Some nice effects is that it reduces thread contention, allows filtering to occur a lot earlier, reduces session wakeups, and simplifies how waking up sessions works.

I’ve done some preliminary artificial testing myself and it has reduced the manager taskprocessor max size by 30-40%. The reason I’m posting on here is to get anyone with heavy AMI usage to give it a try if they can and let me know if they see any difference in general, and for the manager taskprocessor as well.

I do have the ability to get rid of the taskprocessor entirely but I’d like to see the actual impact of this change first with the taskprocessor in use for others.

Does it mean one managers usage cannot be seen by another? That was the first thing I noticed and maybe that’s good but maybe it should have a master manager option that allows them see everything? However that would probably reduce the performance back to the way it was (everyone receives everything)?

No. It doesn’t mean that. There has been no change to visibility of events.

When I command it on one instance the results don’t show on another. I’ll have to recompile the old to see if that is different.

That’s an AMI action, which this code hasn’t touched at all. AMI action responses go to the manager session that did the AMI action.

is it advisable to test on a production server ?

I don’t anticipate issues, but that’s up to you.

i will start tonight need off time to compile

what data points can i get ?
asterisk -rx " core show taskprocessors"|grep manager
stasis/m:manager:core-0000000e 1590891970 145 99302 2700 3000 0 1260378

Processed In Queue Max Depth Low water High water Low time(us) High time(us)
1602323195 145 99302 2700 3000 0 1260378

That taskprocessor output is sufficient.

Processed In Queue Max Depth Low water High water Low time(us) High time(us) patch
4699729189 0 2153074 2700 3000 0 195289 old
1115909 73241 2700 3000 0 37981 new

this is a newly patched system

but is not using the disabledevents filter

just added it is there a way to reset taskprocessors counting or just a restart ?

pbx02*CLI> core reset taskprocessor stasis/m:manager:core-0000000e

There is no reset. They are intended for the duration of the process.

wonder in queue 23591 from where did i get that ?

it cleared up but that is a lot for in queue ?
ahh
Max Depth is the aftermath of in queue ,

ok so there is a queue …

the question is why and is it bothering ?

and to what extent !

Your post previously never had an “in queue” value. Max Depth is the maximum depth the queue ever got.

As for why, events are being published to the manager topic faster than they can be turned into AMI events and queued to applicable manager sessions. As long as the subscription uses a taskprocessor, that’s always going to be a possibility. The alternative is to not use a taskprocessor and instead block every thread when they publish - which I may end up doing in the end, as the tradeoff could be worth it but it’s too early.

Processed In Queue Max Depth Low water High water Low time(us) High time(us)
1602323195 145 99302 2700 3000 0 1260378 OLD
7266676 0 5586 2700 3000 0 102811 NEW
29809718 0 5586 2700 3000 0 632733 NEW
50546980 0 5586 2700 3000 0 632733 NEW

server b

Processed In Queue Max Depth Low water High water Low time(us) High time(us) patch
4699729189 0 2153074 2700 3000 0 195289 old
1115909 0 73241 2700 3000 0 37981 new
180451813 0 804417 2700 3000 0 92289 new

Early days, but seems better. Reduced the amount of time it takes, which lowers queue max depth.

Processed In Queue Max Depth Low water High water Low time(us) High time(us) time DATE
1602323195 145 99302 2700 3000 0 1260378 OLD
7266676 0 5586 2700 3000 0 102811 NEW
29809718 0 5586 2700 3000 0 632733 NEW
50546980 0 5586 2700 3000 0 632733 NEW
97377684 0 5878 2700 3000 0 632738 16:09 16/06/2026

server B

Processed In Queue Max Depth Low water High water Low time(us) High time(us) time DATE
4699729189 0 2153074 2700 3000 0 195289 old
1115909 0 73241 2700 3000 0 37981 new
180451813 0 804417 2700 3000 0 92289 new
279729936 0 804417 2700 3000 0 92289 new
394083919 0 804417 2700 3000 0 105288 16:10 16/06/2026

It continues to behave as I wanted.

server bbx02 is flawed statistics

as added the disabled list only after the patch so the before (controller) and after are not apples to apples

the reason i added it because if not the number at startup where bigger …

01

Processed In Queue Max Depth Low water High water Low time(us) High time(us) time DATE Max Depth High time(us) Processed
1602323195 145 99302 2700 3000 0 1260378 controller
7266676 0 5586 2700 3000 0 102811 NEW 5.63% 8% 0%
29809718 0 5586 2700 3000 0 632733 NEW 5.63% 50% 2%
50546980 0 5586 2700 3000 0 632733 NEW 5.63% 50% 3%
97377684 0 5878 2700 3000 0 632738 16:09 16/06/2026 5.92% 50% 6%
135886100 0 5878 2700 3000 0 632738 12:53 17/06/2026 5.92% 50% 8%
150416944 0 5878 2700 3000 0 632738 17:53 17/06/2026 5.92% 50% 9%

pbx02 also restarted the server yesterday

Processed In Queue Max Depth Low water High water Low time(us) High time(us) time DATE Max Depth High time(us) Processed
4699729189 0 2153074 2700 3000 0 195289 old
1115909 0 73241 2700 3000 0 37981 new 3.40% 19% 0%
180451813 0 804417 2700 3000 0 92289 new 37.36% 47% 4%
279729936 0 804417 2700 3000 0 92289 new 37.36% 47% 6%
394083919 0 804417 2700 3000 0 105288 16:10 16/06/2026 37.36% 54% 8%
115373114 0 75221 2700 3000 0 75221 14:24 17/06/2026 3.49% 39% 2% server restart
151086329 0 75221 2700 3000 0 91751 17:10 17/06/2026 3.49% 47% 3%

05 is very interesting max depth is higher but i only have a controller of a half a day

Processed In Queue Max Depth Low water High water Low time(us) High time(us) time DATE Max Depth High time(us) Processed
543481706 0 103226 2700 3000 0 196395 controller
179783197 0 42485 2700 3000 0 119270 new 41.16% 61% 33%
245035509 629 42485 2700 3000 0 119270 12:45 17/06/2026 41.16% 61% 45%
336557040 0 119308 2700 3000 0 119760 13:54 17/06/2026 115.58% 61% 62%
497986162 130756 2700 3000 0 119760 16:44 17/06/2026 126.67% 61% 92%
555086186 0 130756 2700 3000 0 119760 17:56 17/06/2026 126.67% 61% 102%

look i do not have a logger for extension connected channels and average .

but i am hitting new highs with less of load average

image

something is definitely being doon better causing the load to stay low .

even though the actual max depth and time is going up

pbx5

Processed In Queue Max Depth Low water High water Low time(us) High time(us) time DATE Max Depth High time(us) Processed
543481706 0 103226 2700 3000 0 196395 controller
179783197 0 42485 2700 3000 0 119270 new 41.16% 61% 33%
245035509 629 42485 2700 3000 0 119270 12:45 17/06/2026 41.16% 61% 45%
336557040 0 119308 2700 3000 0 119760 13:54 17/06/2026 115.58% 61% 62%
497986162 130756 2700 3000 0 119760 16:44 17/06/2026 126.67% 61% 92%
555086186 0 130756 2700 3000 0 119760 17:56 17/06/2026 126.67% 61% 102%
1186879383 0 130756 2700 3000 0 176554 15:34 18/06/2026 126.67% 90% 218%
2268749048 305419 2700 3000 0 3410627 12:16 12:16 295.87% 1737% 417%

pbx01

Processed In Queue Max Depth Low water High water Low time(us) High time(us) time DATE Max Depth High time(us) Processed
1602323195 145 99302 2700 3000 0 1260378 controller
7266676 0 5586 2700 3000 0 102811 NEW 5.63% 8% 0%
29809718 0 5586 2700 3000 0 632733 NEW 5.63% 50% 2%
50546980 0 5586 2700 3000 0 632733 NEW 5.63% 50% 3%
97377684 0 5878 2700 3000 0 632738 16:09 16/06/2026 5.92% 50% 6%
135886100 0 5878 2700 3000 0 632738 12:53 17/06/2026 5.92% 50% 8%
150416944 0 5878 2700 3000 0 632738 17:53 17/06/2026 5.92% 50% 9%
253093156 0 8462 2700 3000 0 821842 12:22 21/06/2026 8.52% 65% 16%

There’s nothing inherently in it that would block for long periods of time, aside from the thread itself just being starved - which can happen. Being more efficient though normal operation would probably use less CPU, and when it does hit high water mark it would be able to clear it faster.