
Dynamic memory allocation is a powerful feature in embedded systems, but it introduces risks such as fragmentation, leaks, and allocation failures. In Zephyr RTOS, dynamic memory allocation failures can lead to system instability, crashes, or undefined behavior. This article explores the root causes of such failures, provides debugging techniques, and offers practical solutions to ensure reliable memory usage in your Zephyr applications.
Initial Contiguous System Heap (128 Bytes Free):+---------------------------------------------------------------+| 128 Bytes Free Pool |+---------------------------------------------------------------+Four Interleaved Allocations (A=32B, B=32B, C=32B, D=32B):+-----------+-----------+-----------+-----------+| Chunk A | Chunk B | Chunk C | Chunk D || (32B) | (32B) | (32B) | (32B) || [Alloc'd] | [Alloc'd] | [Alloc'd] | [Alloc'd] |+-----------+-----------+-----------+-----------+External Fragmentation: Chunks A and C Freed (B and D remain live):+-----------+-----------+-----------+-----------+| Chunk A | Chunk B | Chunk C | Chunk D || (32B) | (32B) | (32B) | (32B) || [FREE] | [Alloc'd] | [FREE] | [Alloc'd] |+-----------+-----------+-----------+-----------+Total Free = 64B across two non-adjacent 32B spans.Attempt to Allocate 64 Contiguous Bytes --> FAILS (-ENOMEM):+-----------+-----------+-----------+-----------+| 32B | 32B | 32B | 32B || [FREE] | [BLOCKED] | [FREE] | [BLOCKED] |+-----------+-----------+-----------+-----------+Largest contiguous free span = 32B < 64B requested.Result: k_malloc(64) returns NULL despite 64B total free.
Over time, repeated allocation and deallocation of memory blocks of varying sizes can lead to severe heap fragmentation. In Zephyr RTOS, dynamic memory allocations through k_malloc() and k_heap_alloc() draw from a Cartesian-tree / multi-bucket allocator (sys_heap). Although sys_heap automatically coalesces adjacent free chunks during deallocation, it cannot relocate or compact live memory.
Fragmentation manifests in two forms:
For example, if you allocate four 32-byte blocks (A, B, C, and D) and later free A and C while retaining B and D, the heap contains 64 bytes of total unallocated space split across two non-adjacent 32-byte spans (between B and D, and before B). If an incoming task requests a 64-byte contiguous buffer, the allocator traverses its free lists and discovers that no single chunk can satisfy the request, returning NULL (-ENOMEM) despite sufficient aggregate free RAM.
A memory leak occurs when dynamically allocated memory is never returned to the heap, steadily diminishing the available pool until allocations fail. In embedded Zephyr applications, memory leaks frequently stem from:
k_free().k_msgq) or FIFO channels (k_fifo) where the consumer assumes static ownership or fails to pair deallocation on packet rejection.Because memory leaks consume RAM gradually, they rarely trigger immediate crashes. Instead, the application degrades silently over hours or days until a critical allocation fails, precipitating hard faults or watchdogs.
If the configured heap size is smaller than the peak concurrent allocation demand of the application, exhaustion occurs even in the absence of leaks or fragmentation. The global system heap in Zephyr is sized using CONFIG_HEAP_MEM_POOL_SIZE (specified in bytes in prj.conf).
Underestimating peak memory demand is particularly common in networking (Zephyr Bluetooth LE, OpenThread, or TCP/IP socket buffers) and sensor hub processing pipelines. When multiple high-priority events burst simultaneously, temporary buffers exhaust the available pool. Sizing the heap requires empirical high-water mark profiling rather than nominal steady-state estimation.
Misusing Zephyr’s memory management primitives creates subtle bugs, memory corruption, and allocation panics:
k_malloc(), k_calloc(), k_heap_alloc()) return NULL upon failure. Failing to validate the returned pointer before dereference induces immediate HardFaults or MPU memory management exceptions.k_malloc() executes with K_NO_WAIT and technically does not sleep, allocating from sys_heap within an Interrupt Service Routine (ISR) is a dangerous anti-pattern. Heap traversal operates under a spinlock; executing variable-duration chunk searches inside an ISR introduces unpredictable latency jitter and violates real-time determinism. Furthermore, attempting to call k_heap_alloc() with any blocking timeout (timeout > 0) from an ISR triggers a fatal kernel assertion failure (z_is_in_isr()).k_mem_slab_alloc() into k_free(), or conversely returning a k_malloc() pointer with k_mem_slab_free(), corrupts allocator bookkeeping headers and leads to catastrophic heap breakdown.k_free() multiple times on the same pointer corrupts the free-list chunk pointers inside sys_heap, causing subsequent allocation loops or kernel oops.Adjust CONFIG_HEAP_MEM_POOL_SIZE in your project’s prj.conf to accommodate peak buffer utilization:
CONFIG_HEAP_MEM_POOL_SIZE=32768
While increasing heap capacity resolves immediate capacity constraints, it consumes dedicated SRAM that cannot be reclaimed by thread stacks or static bss segments. If root causes like memory leaks or external fragmentation remain unaddressed, increasing the heap merely defers the eventual failure.
Zephyr provides robust instrumentation tools to monitor allocation dynamics and pinpoint uncollected memory:
CONFIG_SYS_HEAP_RUNTIME_STATS): Enables tracking of current allocated bytes, total free bytes, and historical high-water marks (max_allocated_bytes).CONFIG_SHELL=y and CONFIG_KERNEL_SHELL=y): Exposes the live kernel heap command over serial UART or RTT, allowing engineers to query heap health in real time without external debuggers.CONFIG_SYS_HEAP_VALIDATE): Instruments sys_heap routines to perform integrity audits on boundary tags and chunk pointers during every allocation and release.k_mem_slab)For high-frequency, repetitive allocations (such as network packets, telemetry frames, or sensor sample queues), dynamic heaps should be replaced with Zephyr Memory Slabs.
The k_mem_slab kernel primitive manages a pre-allocated pool of fixed-size memory blocks using an internal free list:
k_mem_slab_alloc()) and release (k_mem_slab_free()) operate in deterministic, constant time, making them safe and predictable even under heavy system load.K_NO_WAIT without risking heap spinlock latency spikes.In multi-resource initialization, partial allocations must be rigorously unwound when a subsequent allocation fails. Adopting the canonical structured cleanup pattern (goto fail or single-point-of-exit) ensures no orphaned memory blocks remain allocated when errors occur.
This example demonstrates how to allocate multiple dynamic buffers and employ a structured cleanup label to prevent memory leaks if a secondary allocation fails:
#include <zephyr/kernel.h>#include <zephyr/logging/log.h>#include <string.h>#include <errno.h>LOG_MODULE_REGISTER(mem_mgmt, LOG_LEVEL_INF);int construct_network_payload(size_t header_sz, size_t payload_sz,uint8_t **out_hdr, uint8_t **out_payload) {uint8_t *header_buf = NULL;uint8_t *payload_buf = NULL;/* Primary buffer allocation */header_buf = (uint8_t *)k_malloc(header_sz);if (header_buf == NULL) {LOG_ERR("Failed to allocate header buffer (%zu bytes)", header_sz);return -ENOMEM;}/* Secondary buffer allocation */payload_buf = (uint8_t *)k_malloc(payload_sz);if (payload_buf == NULL) {LOG_ERR("Failed to allocate payload buffer (%zu bytes)", payload_sz);goto fail_payload;}/* Initialise buffers */memset(header_buf, 0x00, header_sz);memset(payload_buf, 0x00, payload_sz);/* Transfer ownership to caller */*out_hdr = header_buf;*out_payload = payload_buf;return 0;fail_payload:k_free(header_buf);return -ENOMEM;}
k_mem_slabThis example demonstrates setting up and consuming a zero-fragmentation memory slab for uniform packet buffers:
#include <zephyr/kernel.h>#include <zephyr/sys/printk.h>#define PACKET_BLOCK_SIZE 128#define PACKET_BLOCK_COUNT 16#define PACKET_ALIGNMENT 4/* Statically define slab: 16 blocks of 128 bytes each, 4-byte aligned */K_MEM_SLAB_DEFINE(packet_slab, PACKET_BLOCK_SIZE, PACKET_BLOCK_COUNT, PACKET_ALIGNMENT);int process_telemetry_frame(const uint8_t *src_data, size_t len) {void *slab_block = NULL;if (len > PACKET_BLOCK_SIZE) {return -EINVAL;}/* O(1) allocation with non-blocking timeout */int status = k_mem_slab_alloc(&packet_slab, &slab_block, K_NO_WAIT);if (status != 0) {printk("Packet slab exhausted (status: %d)\n", status);return -ENOMEM;}memcpy(slab_block, src_data, len);/* Process packet ... *//* Release block back to slab */k_mem_slab_free(&packet_slab, slab_block);return 0;}
Add the following lines to your project configuration to enable heap debugging and monitoring:
# Sizing the global kernel system heapCONFIG_HEAP_MEM_POOL_SIZE=32768# Enable runtime heap statistics collectionCONFIG_SYS_HEAP_RUNTIME_STATS=y# Enable heap integrity boundary auditsCONFIG_SYS_HEAP_VALIDATE=y# Enable Zephyr Shell subsystem and Kernel Shell moduleCONFIG_SHELL=yCONFIG_KERNEL_SHELL=y
This configuration sets up a 32 KB system heap, activates real-time allocation telemetry, enables integrity boundary audits, and exposes the interactive kernel heap shell command.
While distinct from heap allocations, stack overflows often corrupt nearby heap boundaries or occur when developers replace failed heap calls with excessive automatic stack variables:
# Enable hardware MPU/MMU stack overflow protectionCONFIG_HW_STACK_PROTECTION=y# Enable runtime thread stack tracking and high-water mark loggingCONFIG_THREAD_STACK_INFO=yCONFIG_INIT_STACKS=yCONFIG_THREAD_ANALYZER=yCONFIG_THREAD_ANALYZER_AUTO=yCONFIG_THREAD_ANALYZER_AUTO_INTERVAL=10
sys_heap_runtime_stats_getFor headless or automated monitoring, applications can query heap statistics directly:
#include <zephyr/kernel.h>#include <zephyr/sys/sys_heap.h>#include <zephyr/logging/log.h>LOG_MODULE_REGISTER(heap_monitor, LOG_LEVEL_INF);#if defined(CONFIG_SYS_HEAP_RUNTIME_STATS)void log_system_heap_health(struct sys_heap *heap) {struct sys_memory_stats stats;if (sys_heap_runtime_stats_get(heap, &stats) == 0) {LOG_INF("Heap Status: Free=%zu B, Allocated=%zu B, Peak Usage=%zu B",stats.free_bytes, stats.allocated_bytes, stats.max_allocated_bytes);} else {LOG_WRN("Unable to read heap runtime statistics");}}#endif
With CONFIG_SHELL=y, CONFIG_KERNEL_SHELL=y, and CONFIG_SYS_HEAP_RUNTIME_STATS=y compiled in, connect to the device console and execute:
uart:~$ kernel heapfree: 24192 Ballocated: 8576 Bmax alloc: 14336 B
If allocated steadily increases over repeated operational cycles while free fails to return to baseline, a memory leak is actively depleting RAM.
Implement a firmware soak test that allocates and deallocates buffers with varying lifespans to expose fragmentation boundaries:
#include <zephyr/kernel.h>#include <zephyr/sys/printk.h>#define STRESS_ITERATIONS 100#define MAX_CHUNKS 32void run_heap_stress_test(void) {void *allocated_chunks[MAX_CHUNKS] = {NULL};printk("Starting heap fragmentation stress test...\n");for (int iter = 0; iter < STRESS_ITERATIONS; iter++) {int idx = iter % MAX_CHUNKS;/* Free prior chunk if present */if (allocated_chunks[idx] != NULL) {k_free(allocated_chunks[idx]);allocated_chunks[idx] = NULL;}/* Allocate dynamic variable-sized chunk */size_t req_sz = ((iter * 17) % 256) + 16;allocated_chunks[idx] = k_malloc(req_sz);if (allocated_chunks[idx] == NULL) {printk("STRESS TEST FAILURE: Allocation failed at iter %d for %zu bytes\n",iter, req_sz);break;}}/* Final cleanup */for (int i = 0; i < MAX_CHUNKS; i++) {if (allocated_chunks[i] != NULL) {k_free(allocated_chunks[i]);}}printk("Stress test complete.\n");}
When CONFIG_SYS_HEAP_VALIDATE=y is set, invoke sys_heap_validate() periodically across unit tests to assert that chunk headers, guard words, and free list pointers remain uncorrupted by buffer overruns.
Deploy firmware builds under full peripheral stress (BLE advertisements, sensor DMA transfers, network socket transactions) for 48 to 72 hours. Validate that the high-water mark (max alloc) plateaus beneath safe operating limits and does not crawl toward heap exhaustion.
Dynamic memory allocation failures in Zephyr RTOS typically result from external heap fragmentation, unhandled error leaks, undersized heap configuration, or misuse of allocators in interrupt contexts.
Engineers can resolve and prevent these failures by:
k_mem_slab memory pools.kernel heap shell command and CONFIG_SYS_HEAP_RUNTIME_STATS.CONFIG_SYS_HEAP_VALIDATE and hardware MPU stack protection (CONFIG_HW_STACK_PROTECTION).k_mem_slab)“. https://docs.zephyrproject.org/latest/kernel/services/data_passing/memory_slabs.htmlQuick Links
Legal Stuff




