The log is readable, but I think it is one layer downstream of what would explain it. Four numbers would change how I read it.
Answer rate, or drop per dial. Arena count is not driven by peak concurrency, it is driven by thread churn. 763 concurrent at a 20% answer rate means most originates die in six or eight seconds on no-answer, busy or SIT, so you are creating and burning something like 90 threads a second, each reaching for an arena. 763 concurrent connected calls averaging three minutes is the same peak number and would not produce 256 arenas. If your answer rate is low, churn built that ceiling and that is where the fix lives.
Block size and how dials are paced out. Five hundred originates fired at once spin up arenas that the same volume staggered would not, because a new arena gets created when a thread collides with a locked one. Same throughput, different arena count, purely from burst shape.
Endpoints, and whether there is transcoding or recording in that path. This is the one I want most. 26.75 GiB over 241,357 chunks averages about 116 KiB, and nothing in the plain originate path allocates that. The mean is hiding a bimodal distribution: 240k small chunks you would expect, plus a few thousand very large ones holding most of the gigabytes, and the large ones are what pin your heaps. If you are transcoding or writing recordings, those are your codec and file buffers and you already have your answer. If it is passthrough G.711 with nothing recorded, they are something else and worth chasing. malloc_info(0, stream) gives you the per-arena bin histogram as XML if you want it confirmed.
Whether 763 is a configured cap or just where pacing landed.
Then the measurement I would take before the A/B: does RSS climb campaign over campaign, or plateau? Those 188 orphaned arenas are not lost. glibc parks them on a free list and hands them to the next thread that asks, so your next campaign should reuse them rather than allocate 188 more. If you sit at 41 peak / 27 idle across five campaigns, that is a high-water mark parked rather than a leak, and MALLOC_ARENA_MAX buys you nothing malloc_trim would not. If it climbs campaign over campaign, that is a different and much more serious problem.
On the tuning itself, once the above is known.
malloc_trim(0) at campaign end is the cheaper experiment and I would run it first. Current glibc walks every arena, not just main. If it hands back most of that 26.75 GiB you are done without a config change.
MALLOC_ARENA_MAX=8 against 763 concurrent originate threads is about 95 threads per arena. glibc degrades gracefully, arena_get_retry moves on rather than blocking, but on a dialer doing 90 setups a second you will see it. Watch call setup rate and AMI round trip next to RSS, or you will trade a memory number nobody is paying for against a throughput number somebody is.
On bounding the model: the per-channel PBX thread is fundamental to Asterisk and is not going anywhere. The async AMI originate thread is the extra one on top of it, and you cannot pool that from config. You can bound it from your side by capping in-flight originates in the dialer. If the answer to the first question is a low answer rate, that cap is also the highest-leverage change available to you, because it takes churn down at the same time as peak.
Follow-up.
If you like, DM me the additional logs and I’ll have someone dig deeper on this. Good troubleshooting practice for us so would be happy to help out.