High RSS / OOM after heavy AMI Originate campaigns — looks like glibc arena fragmentation, not a leak

We run big outbound campaigns via concurrent AMI Originate (peaked around 763 channels / 1200 threads, RSS ~41 GiB). After the campaign drains and threads exit, RSS stays around ~27 GiB with only 68 threads left.

Grabbed a core in that post-load state and walked glibc’s malloc arenas with GDB:

256 arenas (glibc's default ceiling = 8 × 32 CPUs, no MALLOC_ARENA_MAX set)
188 arenas with attached_threads == 0, holding 18.64 GiB
26.98 of 27.21 GiB of arena memory already free per malloc's own metadata
  - only 0.24 GiB in top chunks
  - 26.75 GiB scattered across normal bins (241357 chunks)

So this doesn’t look like live Asterisk objects — almost all the retained memory is already “freed” from malloc’s point of view, just fragmented across arenas and never handed back to the OS.

We’re considering testing MALLOC_ARENA_MAX=8 to see if capping arena growth prevents this. Curious if anyone else running heavy Originate/dialer workloads has tuned around this, tried malloc_trim(), or has thoughts on whether the Originate thread-per-call model could be bounded/pooled.

Happy to share the core, full GDB transcript, thread dumps, or RSS timeline if useful.

The log is readable, but I think it is one layer downstream of what would explain it. Four numbers would change how I read it.

Answer rate, or drop per dial. Arena count is not driven by peak concurrency, it is driven by thread churn. 763 concurrent at a 20% answer rate means most originates die in six or eight seconds on no-answer, busy or SIT, so you are creating and burning something like 90 threads a second, each reaching for an arena. 763 concurrent connected calls averaging three minutes is the same peak number and would not produce 256 arenas. If your answer rate is low, churn built that ceiling and that is where the fix lives.

Block size and how dials are paced out. Five hundred originates fired at once spin up arenas that the same volume staggered would not, because a new arena gets created when a thread collides with a locked one. Same throughput, different arena count, purely from burst shape.

Endpoints, and whether there is transcoding or recording in that path. This is the one I want most. 26.75 GiB over 241,357 chunks averages about 116 KiB, and nothing in the plain originate path allocates that. The mean is hiding a bimodal distribution: 240k small chunks you would expect, plus a few thousand very large ones holding most of the gigabytes, and the large ones are what pin your heaps. If you are transcoding or writing recordings, those are your codec and file buffers and you already have your answer. If it is passthrough G.711 with nothing recorded, they are something else and worth chasing. malloc_info(0, stream) gives you the per-arena bin histogram as XML if you want it confirmed.

Whether 763 is a configured cap or just where pacing landed.

Then the measurement I would take before the A/B: does RSS climb campaign over campaign, or plateau? Those 188 orphaned arenas are not lost. glibc parks them on a free list and hands them to the next thread that asks, so your next campaign should reuse them rather than allocate 188 more. If you sit at 41 peak / 27 idle across five campaigns, that is a high-water mark parked rather than a leak, and MALLOC_ARENA_MAX buys you nothing malloc_trim would not. If it climbs campaign over campaign, that is a different and much more serious problem.

On the tuning itself, once the above is known.

malloc_trim(0) at campaign end is the cheaper experiment and I would run it first. Current glibc walks every arena, not just main. If it hands back most of that 26.75 GiB you are done without a config change.

MALLOC_ARENA_MAX=8 against 763 concurrent originate threads is about 95 threads per arena. glibc degrades gracefully, arena_get_retry moves on rather than blocking, but on a dialer doing 90 setups a second you will see it. Watch call setup rate and AMI round trip next to RSS, or you will trade a memory number nobody is paying for against a throughput number somebody is.

On bounding the model: the per-channel PBX thread is fundamental to Asterisk and is not going anywhere. The async AMI originate thread is the extra one on top of it, and you cannot pool that from config. You can bound it from your side by capping in-flight originates in the dialer. If the answer to the first question is a low answer rate, that cap is also the highest-leverage change available to you, because it takes churn down at the same time as peak.

Follow-up.
If you like, DM me the additional logs and I’ll have someone dig deeper on this. Good troubleshooting practice for us so would be happy to help out.

Thanks — the pacing/churn and plateau-vs-growth tests are useful.

One thing I should also mention is that I did some broader research after finding this. The same glibc arena-retention pattern has been reported in several unrelated high-concurrency applications (Presto, Ruby, Java, Go/CGO, Python services, etc.), and limiting MALLOC_ARENA_MAX is a commonly documented mitigation.

My core also already shows:

27.21 GiB total arena system_mem
26.98 GiB already free
241,357 free chunks
188 arenas with attached_threads == 0
256 arenas on a 32-core host

So I don’t think we can infer from the ~116 KiB average free-chunk size alone that codec/recording buffers are responsible.

Agreed on the average, it settles nothing by itself. That was the point of flagging it: it is the reason to pull the histogram, not evidence on its own. malloc_info(0, stream) gives the per-bin breakdown per arena and either shows a few thousand chunks sitting in the large bins or it does not. Worth noting you need the live process at the same point in the cycle for that, not the core.

On the cross-application reports, I would be careful transplanting them. Presto, the JVM, Ruby and Go/CGO are steady-state pooled workloads: a fixed set of long-lived threads, arenas created once and reused for the life of the process. Capping MALLOC_ARENA_MAX helps there because the arena count really is the whole problem.

Yours is a churn workload with the same symptom and a different mechanism. 763 concurrent originates at a low answer rate means threads created and destroyed continuously, and arenas orphaned at their high-water mark, rather than a fixed set that grew. That difference is exactly why plateau-versus-growth comes first. In the pooled cases the memory genuinely never returns. In a churn workload glibc hands an orphaned arena to the next thread that asks for one, so campaign two may cost you nothing at all, and if it plateaus then ARENA_MAX is buying you contention on a box doing ninety setups a second to fix a number nobody was paying.

Two things settle it and neither needs a core: RSS at idle after campaign three next to campaign one, and your answer rate.

After campaign load subsided, Asterisk retained ~20.6 GB RSS. A direct in-process malloc_trim(0) reduced RSS to ~1.07 GB, reclaiming ~18.65 GiB in ~3.0 seconds without restarting Asterisk.

-this is a ‘real human’ response-
This answers the leak question. If a few seconds of trim hands that much back, none of it was ever lost. It was free the whole time and the allocator just wasn’t giving it up.

It’s not a fix tho, run the next campaign and you’ll be right back where you started. Nothing really in the way it allocates changed, you just emptied the bucket.

You did get a clean floor, so a trim after the first campaign and write the number down then trim again after the third and compare. If you notice it creeping up something is leaking and you should be able to see it. But, if the numbers land in the same place both times, nothing is leaking and a capped arena count along with the trim between runs is the whole story.

It’s not a loss to run the arena cap on its own. Fewer arenas would probably mean less of this noise but with the thread count you’re showing you may be just trading for contention.

very interesting i updated asterisk to 20.21.0

which has a AMI improvement [1] from JCOLP
(using AMI vey heavy)

and also i set taskprocessors from 4 to 25 and the memory cleared by it self

[1] manager: Move away from shared linked list for events. by jcolp · Pull Request #1984 · asterisk/asterisk · GitHub

Good result, and it explains the whole thing. Nothing was leaking. It was how much got allocated and how long it sat there, and you cut both, so the allocator had nothing left to hold.

One note. You changed two things at once, so if it ever comes back you won’t know which one was carrying it. Write down what you’re on now while you still remember.

Thanks for posting the fix. Most people don’t come back.