Files
SkinbaseNova/docs/optimization-m6-php-fpm-memory.md
T
klevze 8a80aae21e Ship production optimization M1-M12.5A: queues, metrics, HTTP observability, and vector search reliability.
Keep similar-ai from tripping the global circuit on a lone URL 502, clamp Qdrant search to 100, and add Server-Timing plus slow-request logging. Studio shared props, Academy S3 exists caching, heat chunking, and Redis/scheduler hygiene stay in this rollout.
2026-08-25 07:58:47 +02:00

12 KiB
Raw Blame History

M6 — PHP-FPM, Memory, Swap & Worker Capacity

STATUS: COMPLETE (read-only; no production changes)
WHEN: 2026-08-23 15:29–15:32 Europe/Ljubljana (UTC 13:29)
HOST: server3, Debian 13, kernel 6.12.96, 8 cores, uptime 24d 18h

Nothing was restarted, tuned, flushed, or edited on server3.


Verdict

Swap:              KEEP 2 GiB. Nearly full, but NOT thrashing (si/so=0, PSI memory=0).
PHP-FPM max_children=14:  KEEP. Peak has used all 14; listen queue never built.
Do not raise FPM/Horizon on this host until RAM is freed or the machine is larger.
InnoDB 10 GB pool: KEEP for 30-day snapshot growth (~7.5 GB table). Currently ~24% filled.
OPcache 256 MB:    KEEP (runtime used unknown; config is not tight on paper).
SSR:               ~116 MB RSS; not a memory problem. Frequent SIGTERM restarts are operator/supervisor, not OOM.
Redis:             879 MB used / 2 GB max; RSS 167 MB — likely swapped. 0 evictions.

Largest memory consumers are MySQL (8.1 GB RSS), vision Docker/Python (~2.5 GB), ClamAV (1.1 GB) — not Skinbase FPM.


1. Server RAM / swap

Metric Value
MemTotal 23 GiB (24,615,736 kB)
Mem used 15 GiB
MemAvailable 8.2 GiB
Buffers+Cached 7.2 GiB
AnonPages 13.4 GiB
SwapTotal 2.0 GiB
Swap used 2.0 GiB (37 MiB free)
SwapCached 774 MiB
Committed_AS 30.1 GiB
vm.swappiness 10
vm.overcommit_memory 0
Load 5.09 / 7.07 / 8.33 (8 CPUs)
PSI memory some/full avg10=0.00
PSI cpu some avg10=4.7, avg60=14.5
vmstat si/so (1s×5) 0 / 0
Processes 243

Swap interpretation: cold / previously used pages sitting in swap, with 774 MiB SwapCached. Not active thrashing. Do not disable swap. Do not resize without a RAM plan. swappiness=10 is already conservative.


2. Process memory (RSS, measured)

Process RSS Notes
mysqld 8087 MB InnoDB pool configured 10 GB
uvicorn/python (vision, port 8000) 1310 + 542 + 238 + 197 + 103 + 54 MB Docker vision stack
clamd 1107 MB
qdrant 262 MB vision
meilisearch 395 MB Skinbase search
crowdsec 296 MB
redis-server 171 MB RSS / 879 MB used_memory swapped
netdata 180 MB
gitea 156 MB
php-fpm master 48 MB
php-fpm pool skinbase (×6) 88–110 MB each see §4
php-fpm pool www (×2) 13 MB each unused by Skinbase vhost
horizon master 104 MB
horizon supervisors (×3) ~103 MB each
horizon workers (×6 now) 103–125 MB --memory=128
node SSR 116 MB uptime minutes (restarted)
reverb 106 MB
nginx workers 10–58 MB one “shutting down” 15d
php8.4-fpm cgroup current 746 MB, peak 1575 MB all pools

3. PHP-FPM Skinbase pool (php8.4-fpm-skinbase.sock)

From /etc/php/8.4/fpm/pool.d/skinbase.conf:

Setting Production value
pm dynamic
pm.max_children 14
pm.start_servers 4
pm.min_spare_servers 2
pm.max_spare_servers 6
pm.process_idle_timeout unset (N/A for dynamic)
pm.max_requests 500
request_terminate_timeout 120s
request_slowlog_timeout 10s
slowlog /var/log/php8.4-fpm-skinbase-slow.log
pm.status_path /fpm-status (localhost)
ping.path /fpm-ping
listen.backlog unset → PHP default 511
memory_limit 256M (pool admin)
max_execution_time 60

Separate www pool: pm.max_children=5 on /run/php/php8.4-fpm.sock.


4. FPM status (localhost /fpm-status, pool up ~28.7h since 2026-08-22 10:50)

Field Value
accepted conn 214,108
listen queue 0
max listen queue 0
listen queue len 0
idle processes 5
active processes 1
total processes 6
max active processes 14
max children reached 10
slow requests (since start) 2
memory peak (FPM counter) 60 MB

Not currently exhausting workers. Peak has used all 14 children. No socket backlog was recorded, so nginx did not sit behind a full listen queue — capacity is tight at peak, idle most of the time.

Worker RSS (KB): 87712, 89996, 106444, 109308, 109476, 109480.

Stat RSS
min 86 MB
median 108 MB
mean 102 MB
max 110 MB
PSS not readable (permission)
oldest worker ~7.5 min (max_requests=500 recycling)

memory_limit 256M is the ceiling, not typical RSS. Do not size max_children from 256M × 14.

Safe children from measured RSS:

typical_rss ≈ 110 MB
14 × 110 MB ≈ 1.54 GB   (matches systemd MemoryPeak 1.58 GB for all FPM)
14 × 256 MB ≈ 3.58 GB   worst-case if every worker hits the PHP limit

Raising max_children on a host with swap already full and AnonPages 13.4 GiB is not justified. Keep 14.


5. OPcache (config; FPM runtime stats not exposed)

Setting Value
enable 1
memory_consumption 256 MB
interned_strings_buffer 32 MB
max_accelerated_files 50,000
validate_timestamps 1
revalidate_freq 2 s
JIT disable / 0

CLI cannot read the FPM cache. 50k file slots and 256 MB are large for this Laravel app. Do not increase. Optional later (MEDIUM, not memory): validate_timestamps=0 after atomic releases.


6. Horizon (Supervisor skinbase-horizon)

Production env (config/horizon.php on the server):

Supervisor queues maxProcesses timeout tries memory
supervisor-default search, default 5 960 1 128
supervisor-messaging broadcasts, notifications 3 90 1 128
supervisor-mail mail 2 90 5 128

Measured now: 1 master + 3 supervisors + 6 workers ≈ 1.15 GB RSS.

Peak if all maxProcesses spawn: 1+3+5+3+2 = 14 PHP procs × ~110 MB ≈ 1.5 GB.

Mail isolation is intact (tries=5, own supervisor). Do not merge queues.

Redis lists (prefix as used by Laravel):

List LLEN
queues:default 796
queues:mail 0
queues:search 0
queues:broadcasts 0
queues:notifications 0

HIGH: default queue depth 796 — worker throughput, not FPM RAM. Investigate job mix in a later milestone; do not add Horizon processes until RAM is free (each extra worker ≈ 110 MB and --memory=128 is already near RSS 125 MB on default).

Horizon last supervisor restart: 2026-08-23 14:36 (clean exit 0 / SIGTERM wait), not OOM.


7. SSR

  • RSS 116 MB, Node /opt/www/virtual/SkinbaseNova/bootstrap/ssr/ssr.js
  • Supervisor restarts today: 14:22, 14:30, 14:47, 15:14, 15:16, 15:28 — all SIGTERM “waiting to stop”, then spawn. Not crash loops from memory.
  • Heap not sampled (would require attaching to Node).
  • Not material to host memory pressure.

8. Redis (read-only via Laravel)

Field Value
used_memory 879 MB (peak 917 MB)
used_memory_rss 167 MB
maxmemory 2.00 GB
policy allkeys-lru
fragmentation_ratio 0.19
evicted_keys 0
expired_keys 225,783
connected_clients 28
blocked_clients 0
keys 3773 (3026 with TTL)
AOF off
ops/sec ~50
hit/miss 1.46M / 1.39M

frag 0.19 + RSS << used_memory ⇒ Redis pages are in swap. Latency risk under cache bursts. 0 evictions: cache is within 2 GB. Do not flush. Do not raise maxmemory.


9. MySQL / Percona (read-only)

Field Value
innodb_buffer_pool_size 10.00 GB (10 instances)
pages_data / total 156,554 / 655,360 = 23.9%
pages_free 498,726
pool hit 1 − 104,104 / 23,840,832,084 ≈ 99.9996%
max_connections 120
Max_used_connections 26
Threads_connected / running 12 / 3
tmp tables / disk tmp 1,987,237 / 21
innodb_redo_log_capacity 2 GB
Slow_queries 11,997 (uptime 103,663 s)

10 GB pool is oversized for today’s ~2.8 GB schema, but M5 30-day hourly snapshots grow toward ~7.5 GB. Keeping 10 GB avoids refitting later. The table does not need to sit entirely in the pool; hit ratio is already excellent. Do not shrink now; do not grow.


10. Combined memory budget (measured / plausible peak)

Bucket Steady now Plausible peak
OS/page cache (reclaimable) ~7 GB cache shrinks under pressure
MySQL 8.1 GB ~10 GB pool
Vision Docker/Python/Qdrant ~2.7 GB similar
ClamAV 1.1 GB similar
Meilisearch 0.40 GB 0.5 GB
Crowdsec 0.30 GB similar
Redis RSS 0.17 GB 2 GB if unswapped
PHP-FPM Skinbase 0.65 GB (6×110) 1.54 GB (14×110)
PHP-FPM www 0.03 GB 0.06 GB
Horizon 1.15 GB 1.5 GB
SSR + Reverb 0.22 GB 0.3 GB
nginx 0.20 GB 0.25 GB
Netdata/Gitea/journald/docker ~0.5 GB similar
Swap 2 GB in use —

Steady anonymous ~14 GB + 2 GB swap explains the full swap file while 8.2 GB still Available as cache. Peak Skinbase (FPM 14 + Horizon 14) ≈ 3.0 GB, already observed in FPM cgroup peak 1.58 GB.


11. PHP-FPM recommendation

NO CHANGE to pool sizing.

pm = dynamic
pm.max_children = 14          # keep; peak already 14, listen queue 0
pm.start_servers = 4          # keep
pm.min_spare_servers = 2      # keep
pm.max_spare_servers = 6      # keep (matches current total=6)
pm.max_requests = 500         # keep; workers recycle in minutes

Calculation: median RSS 108 MB; 14×108≈1.5 GB. Available RAM is cache, not idle anonymous. Raising children increases swap of Redis/MySQL. Lowering below 14 would clip the measured peak (max active processes=14).


12. Swap recommendation

Option Decision
Keep 2 GiB swap YES
Resize swap NO (overflow is working; host needs RAM or fewer co-tenants)
Change vm.swappiness NO (already 10)
Disable swap NO

13. OOM / crash safety

Check Result
Kernel OOM (journalctl -k / dmesg) no lines visible to this user
PHP-FPM worker memory YES — Allowed memory size of 268435456 Predis StreamConnection.php 2026-08-13 and 2026-08-18
Redis OOM / evictions no (evicted_keys=0)
MySQL no evidence in readable logs
Horizon clean supervisor stop 14:36, not killed
SSR SIGTERM restarts, not OOM

14. Ranked findings

Sev Finding Change now?
HIGH Swap 2 GiB ~full; Redis RSS≪used_memory No sysctl. Free RAM (vision/clam) or larger VM later
HIGH queues:default LLEN=796 Not FPM. Later queue milestone; don’t add Horizon RAM
HIGH Shared host: vision ~2.7 GB + clamd 1.1 GB Out of Skinbase app; ops placement
MEDIUM FPM max_children reached 10×, max active 14 Keep 14; watch listen queue
MEDIUM PHP 256 MB fatals on Predis App/job payload size, not pool size
MEDIUM slowlog 10s, 139k historical lines; status slow=2 since 22 Aug Leave timeout; log rotation ops
LOW nginx worker “shutting down” 15d Recycle nginx in a maintenance window
LOW www pool unused by Skinbase vhost Optional disable later
NO CHANGE OPcache 256/32/50k
NO CHANGE InnoDB 10 GB Needed as snapshots grow
NO CHANGE Horizon isolation / mail supervisor
NO CHANGE SSR RSS

15. Expected benefit of doing nothing to FPM

Avoid swapping MySQL/Redis further. Current request path is not FPM-bound (active=1, listen queue=0).


16. Deployment plan (only if a change is later approved)

M6 recommends no FPM/swap/sysctl/MySQL/Horizon count change.

If ops later moves vision off-box or adds RAM:

  1. Re-measure RSS and /fpm-status.
  2. Then consider pm.max_children using new median RSS, not 256M.
  3. Pool edit + systemctl reload php8.4-fpm (reload, not restart if possible).
  4. Rollback: restore skinbase.conf and reload.

Any pm.* or memory_limit change requires FPM reload. Swap/sysctl/MySQL buffer pool need service-level restarts — not proposed.


17. Risks of increasing max_children without more RAM

More anonymous PHP, more swap of Redis (cache latency) and InnoDB (buffer pool eviction), possible FPM 256 MB fatals still happen per-request.


Production was not modified.