Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
115 changes: 115 additions & 0 deletions docs/LOW_MEMORY_BOARDS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,115 @@
# Running on Low-Memory Boards

Applies to the Pi Zero 2 W (512 MB), Pi 3 / 3B+ (1 GB), and the 1 GB Pi 4.
If your board has 2 GB or more you can skip this document.

## The failure this prevents

The display process is the largest thing on the board. On a 1 GB Pi 3B+ with
around 20 plugins enabled it settles near **600 MB of 905 MB usable**, leaving
under 200 MB of headroom for everything else.

When that headroom runs out, the board does not crash cleanly. `fork()` starts
failing, and because a new process is needed to do almost anything, the
symptoms look nothing like "out of memory":

| What you see | Why |
|---|---|
| SSH accepts the connection then closes it instantly, before any banner | `sshd` forks a session per connection; the fork fails |
| The web UI still responds quickly | Already running, serves from existing threads, forks nothing |
| Ping is perfect, 0% loss | Handled entirely in the kernel |
| The panel is dark | The display process was killed and cannot be respawned |
| The clock is wrong after the next boot | `fake-hwclock`'s periodic save is a scheduled job, and it cannot fork either |

The board looks healthy from the outside and cannot be logged into. Only a
power cycle clears it. If you are here because SSH stopped working, also see
[SSH_UNAVAILABLE_AFTER_INSTALL.md](SSH_UNAVAILABLE_AFTER_INSTALL.md), which
covers the more common cause (AP mode).

## Check your headroom

```bash
free -m
ps -eo rss,comm --sort=-rss | head -5
```

If `MemAvailable` is under ~150 MB while the display is running, you are close
to the edge. To watch it over time:

```bash
watch -n 30 'free -m | head -2'
```

Available memory that falls steadily rather than holding flat means you will
reach the wall; it is a question of when.

## What to do

**1. Enable the memory cgroup controller.** Without it, the `MemoryMax=85%` in
`systemd/ledmatrix.service` is accepted by systemd and silently ignored, so the
service has no ceiling and a runaway takes the whole board down instead of just
restarting. Raspberry Pi firmware disables this controller by default.

`first_time_install.sh` does this for you. To check it took effect:

```bash
grep memory /sys/fs/cgroup/cgroup.controllers
```

If that prints nothing, add `cgroup_enable=memory cgroup_memory=1` to the
kernel command line and reboot. Edit whichever file your image uses —
`/boot/firmware/cmdline.txt` on current Raspberry Pi OS, `/boot/cmdline.txt` on
older layouts (the installer checks the first and falls back to the second).
Everything must stay on a single line.

This changes the failure mode from "the board becomes unreachable" to "the
display service restarts". It is a safety net, not a fix.

**2. Run fewer plugins.** This is the actual remedy. Every enabled plugin costs
memory permanently — its module, its parsed config, and its cached API
responses. On a 512 MB or 1 GB board, keep the enabled set small and prefer
plugins that poll infrequently.

**3. Lower the cache ceiling.** The in-memory cache is sized from total RAM
(150 entries at 1 GB and below, up to 1500 at 8 GB). To go lower still:

```ini
# /etc/systemd/system/ledmatrix.service.d/override.conf
[Service]
Environment=LEDMATRIX_CACHE_MAX_ENTRIES=75
```
Comment thread
coderabbitai[bot] marked this conversation as resolved.

Writing the file does not change the running service. Reload systemd and
restart it:

```bash
sudo systemctl daemon-reload
sudo systemctl restart ledmatrix
```

Fewer entries means more API calls, so lower this only while you are actually
short of memory.

**4. Consider `MemoryHigh`.** `MemoryMax` kills and restarts. `MemoryHigh`
throttles and reclaims instead, which is gentler — but on a board where the
process genuinely wants more than the limit, sustained reclaim can stall the
render loop and show as visible stutter on the panel. Add it only if you prefer
degraded output to a restart:

```ini
[Service]
MemoryHigh=70%
```

## Keep your logs

These images default to volatile journald storage, so every reboot destroys the
logs — including the ones explaining why the board rebooted. `first_time_install.sh`
enables persistent storage capped at 64 MB. To confirm:

```bash
journalctl --list-boots
```

More than one boot listed means logs are surviving reboots. If only one is
listed, journald is still writing to `/run` (tmpfs).
1 change: 1 addition & 0 deletions docs/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,6 +14,7 @@ the one-shot installer. The pages here go deeper.
5. [TROUBLESHOOTING.md](TROUBLESHOOTING.md) — common issues and fixes
6. [SSH_UNAVAILABLE_AFTER_INSTALL.md](SSH_UNAVAILABLE_AFTER_INSTALL.md) — recovering SSH after install
7. [CONFIG_DEBUGGING.md](CONFIG_DEBUGGING.md) — diagnosing config problems
8. [LOW_MEMORY_BOARDS.md](LOW_MEMORY_BOARDS.md) — Pi Zero 2 W / 3B+ / 1GB Pi 4 memory limits

## I want to write a plugin

Expand Down
31 changes: 29 additions & 2 deletions docs/SSH_UNAVAILABLE_AFTER_INSTALL.md
Original file line number Diff line number Diff line change
Expand Up @@ -20,7 +20,22 @@ The installation script:
- Installs and configures `dnsmasq` (DHCP server for AP mode)
- These services can interfere with normal WiFi client mode

### 3. Reboot After Installation
### 3. The Board Ran Out of Memory

On a 512MB or 1GB board, memory exhaustion stops `sshd` being able to fork a
session process. The connection is accepted and then closed immediately, before
any banner:

```text
kex_exchange_identification: Connection closed by remote host
```

The giveaway is that the board is otherwise healthy — ping is clean and the web
UI still responds — but nothing that needs to start a new process works, and
the panel is usually dark. Only a power cycle clears it. See
[LOW_MEMORY_BOARDS.md](LOW_MEMORY_BOARDS.md).
Comment thread
coderabbitai[bot] marked this conversation as resolved.

### 4. Reboot After Installation

If the script reboots the Pi (which it recommends), network services may restart in a different state, potentially triggering AP mode.

Expand Down Expand Up @@ -190,11 +205,23 @@ The web interface allows you to:

## Summary

**SSH becomes unavailable because**:
**SSH becomes unavailable because** — two unrelated causes, and they need
different responses:

*AP mode (most common):*
- WiFi monitor service enables AP mode when WiFi disconnects
- AP mode switches WiFi from client to access point mode
- Pi loses connection to your original network

*Memory exhaustion (low-memory boards):*
- The board runs out of memory, so `sshd` cannot fork a session process
- The connection is accepted and closed before any banner
- Ping still answers and the web UI still responds, so it looks healthy
- The panel is usually dark and the service cannot restart
- **Only a power cycle clears this** — there is no remote recovery, because
every remote route needs a new process
- Prevention and tuning: [LOW_MEMORY_BOARDS.md](LOW_MEMORY_BOARDS.md)

**To regain SSH**:
1. Connect to **LEDMatrix-Setup** AP network (password: `ledmatrix123`)
2. SSH to `192.168.4.1`
Expand Down
88 changes: 88 additions & 0 deletions first_time_install.sh
Original file line number Diff line number Diff line change
Expand Up @@ -1688,6 +1688,94 @@ else
echo "✗ $CMDLINE_FILE not found; skipping isolcpus optimization"
fi

# Enable the memory cgroup controller (idempotent).
# The Pi firmware boots with cgroup_disable=memory, so systemd's MemoryMax= is
# accepted and silently ignored — the display service then has no ceiling, and
# a runaway takes the whole board down (sshd can no longer fork, the panel goes
# dark) rather than just restarting the one service.
if [ "$SKIP_PERF" != "1" ] && [ -f "$CMDLINE_FILE" ]; then
# Both parameters are required for the memory controller, and they can get
# separated -- an image, another tool or a half-applied earlier run can
# leave one without the other. Checking only cgroup_enable=memory would
# report success while MemoryMax= silently does nothing, so each is checked
# and appended independently.
cgroup_missing=""
for cgroup_param in cgroup_enable=memory cgroup_memory=1; do
if ! grep -qw "$cgroup_param" "$CMDLINE_FILE"; then
cgroup_missing="$cgroup_missing $cgroup_param"
fi
done
if [ -z "$cgroup_missing" ]; then
echo "cgroup memory parameters already present in $CMDLINE_FILE"
else
echo "Adding${cgroup_missing} to $CMDLINE_FILE..."
cp "$CMDLINE_FILE" "$CMDLINE_FILE.bak" 2>/dev/null || true
# The kernel command line must stay on one line.
sed -i "1 s|\$|${cgroup_missing}|" "$CMDLINE_FILE"
echo " Takes effect after reboot. Verify with:"
echo " grep memory /sys/fs/cgroup/cgroup.controllers"
fi
Comment thread
coderabbitai[bot] marked this conversation as resolved.
fi

# Persist the journal (idempotent).
# These images default to volatile storage: journald keeps everything in /run
# (tmpfs), so every reboot destroys the logs — including the ones that would
# explain why the board rebooted. Capped so an SD card is not worn out by logs.
# A non-empty /var/log/journal does not prove journald is configured the way
# this needs: the directory survives a switch back to volatile storage, and it
# says nothing about whether a size cap is set. Read the effective
# configuration instead, and only write the keys the user has not set
# themselves so an explicit local limit is preserved.
journald_effective() {
# systemd-analyze merges journald.conf with every drop-in; grep is the
# fallback for images that ship without it.
if command -v systemd-analyze >/dev/null 2>&1 &&
systemd-analyze cat-config systemd/journald.conf >/dev/null 2>&1; then
systemd-analyze cat-config systemd/journald.conf 2>/dev/null
else
cat /etc/systemd/journald.conf /etc/systemd/journald.conf.d/*.conf 2>/dev/null
fi
}
journald_conf="$(journald_effective)"
journald_storage="$(printf '%s\n' "$journald_conf" | grep -E '^[[:space:]]*Storage=' | tail -n1 | cut -d= -f2 | tr -d '[:space:]')"
journald_cap="$(printf '%s\n' "$journald_conf" | grep -E '^[[:space:]]*SystemMaxUse=' | tail -n1 | cut -d= -f2 | tr -d '[:space:]')"

if [ "$journald_storage" = "persistent" ] && [ -n "$journald_cap" ]; then
echo "Persistent journald storage already configured (SystemMaxUse=$journald_cap)"
else
echo "Enabling persistent journald storage..."
mkdir -p /etc/systemd/journald.conf.d
{
echo "# Installed by LEDMatrix first_time_install.sh"
echo "[Journal]"
echo "Storage=persistent"
if [ -n "$journald_cap" ]; then
echo "# SystemMaxUse left to your existing setting ($journald_cap)"
else
# Capped so logs cannot wear out or fill an SD card.
echo "SystemMaxUse=64M"
fi
} > /etc/systemd/journald.conf.d/ledmatrix-persistent.conf
Comment thread
ChuckBuilds marked this conversation as resolved.
mkdir -p /var/log/journal
systemd-tmpfiles --create --prefix /var/log/journal >/dev/null 2>&1 || true
systemctl restart systemd-journald >/dev/null 2>&1 || true

# Drop-ins are applied in lexical order, so a locally added file that sorts
# after ledmatrix-persistent.conf (zz-local.conf and friends) still wins.
# Writing the file is not evidence it took effect -- re-read and say so
# plainly rather than reporting success we cannot confirm.
journald_now="$(journald_effective | grep -E '^[[:space:]]*Storage=' | tail -n1 | cut -d= -f2 | tr -d '[:space:]')"
if [ "$journald_now" = "persistent" ]; then
echo " Persistent journald storage active"
else
echo " WARNING: journald storage is still '${journald_now:-unset}' after"
echo " writing /etc/systemd/journald.conf.d/ledmatrix-persistent.conf."
echo " Another drop-in that sorts later is overriding it. Check:"
echo " systemd-analyze cat-config systemd/journald.conf | grep -n Storage="
echo " Logs will not survive a reboot until that is resolved."
fi
fi

# Ensure dtparam=audio=off in config.txt (idempotent)
if [ "$SKIP_PERF" = "1" ]; then
: # skipped
Expand Down
91 changes: 75 additions & 16 deletions src/cache/memory_cache.py
Original file line number Diff line number Diff line change
Expand Up @@ -4,11 +4,58 @@
Handles in-memory caching with TTL support, size limits, and automatic cleanup.
"""

import os
import time
import threading
import logging
from typing import Dict, Any, Optional

# Historical fixed ceiling, kept as the fallback when RAM cannot be read.
DEFAULT_MAX_SIZE = 1000


def _total_memory_mb() -> Optional[float]:
"""Physical RAM in MB, or None where /proc/meminfo is unavailable."""
try:
with open('/proc/meminfo', 'r', encoding='utf-8') as fh:
for line in fh:
if line.startswith('MemTotal:'):
return int(line.split()[1]) / 1024
except (OSError, ValueError, IndexError):
return None
return None


def default_max_size() -> int:
"""Entry ceiling scaled to this machine's RAM.

One fixed ceiling cannot serve both a 512 MB Pi Zero 2 W and an 8 GB Pi 5.
Entries here are parsed API payloads that routinely run tens of kilobytes
each, so a thousand of them is a comfortable cache on a large board and a
substantial fraction of total RAM on a small one — where the process
competing for that RAM is also driving the panel. Set
LEDMATRIX_CACHE_MAX_ENTRIES to override.
"""
override = os.environ.get('LEDMATRIX_CACHE_MAX_ENTRIES')
if override:
try:
value = int(override)
if value > 0:
return value
except ValueError:
pass

total_mb = _total_memory_mb()
if total_mb is None:
return DEFAULT_MAX_SIZE
if total_mb < 1536: # 512 MB and 1 GB boards
return 150
if total_mb < 3072: # 2 GB
return 400
if total_mb < 6144: # 4 GB
return 800
return 1500 # 8 GB and up


class MemoryCache:
"""Manages in-memory cache with TTL and size limits."""
Expand Down Expand Up @@ -87,6 +134,32 @@ def set(self, key: str, value: Dict[str, Any]) -> None:
with self._lock:
self._cache[key] = value
self._timestamps[key] = time.time()
# Enforce the ceiling here rather than leaving it to the periodic
# cleanup, which only runs every cleanup_interval seconds (300 by
# default). A burst of inserts between two sweeps could otherwise
# take the cache far past _max_size, which is the memory growth this
# limit exists to prevent -- and on a 1GB board that is the
# difference between a bounded cache and an unreachable Pi.
self._evict_over_limit_locked()

def _evict_over_limit_locked(self) -> int:
"""Drop oldest entries until the cache is within _max_size.

Caller must hold self._lock. Returns the number of entries removed.
"""
excess = len(self._cache) - self._max_size
if excess <= 0:
return 0
oldest = sorted(
self._timestamps.items(),
key=lambda item: float(item[1]) if isinstance(item[1], (int, float)) else 0.0
)
removed = 0
for key, _ in oldest[:excess]:
self._cache.pop(key, None)
self._timestamps.pop(key, None)
removed += 1
return removed

def clear(self, key: Optional[str] = None) -> None:
"""
Expand Down Expand Up @@ -143,22 +216,8 @@ def cleanup(self, force: bool = False) -> int:
self._timestamps.pop(key, None)
removed_count += 1

# Enforce size limit by removing oldest entries if cache is too large
if len(self._cache) > self._max_size:
# Sort by timestamp (oldest first)
sorted_entries = sorted(
self._timestamps.items(),
key=lambda x: float(x[1]) if isinstance(x[1], (int, float)) else 0
)

# Remove oldest entries until we're under the limit
excess_count = len(self._cache) - self._max_size
for i in range(excess_count):
if i < len(sorted_entries):
key = sorted_entries[i][0]
self._cache.pop(key, None)
self._timestamps.pop(key, None)
removed_count += 1
# Same ceiling enforcement set() uses, so the two cannot drift.
removed_count += self._evict_over_limit_locked()

self._last_cleanup = current_time

Expand Down
Loading
Loading