Keep similar-ai from tripping the global circuit on a lone URL 502, clamp Qdrant search to 100, and add Server-Timing plus slow-request logging. Studio shared props, Academy S3 exists caching, heat chunking, and Redis/scheduler hygiene stay in this rollout.
12 KiB
M6 — PHP-FPM, Memory, Swap & Worker Capacity
STATUS: COMPLETE (read-only; no production changes)
WHEN: 2026-08-23 15:29–15:32 Europe/Ljubljana (UTC 13:29)
HOST: server3, Debian 13, kernel 6.12.96, 8 cores, uptime 24d 18h
Nothing was restarted, tuned, flushed, or edited on server3.
Verdict
Swap: KEEP 2 GiB. Nearly full, but NOT thrashing (si/so=0, PSI memory=0).
PHP-FPM max_children=14: KEEP. Peak has used all 14; listen queue never built.
Do not raise FPM/Horizon on this host until RAM is freed or the machine is larger.
InnoDB 10 GB pool: KEEP for 30-day snapshot growth (~7.5 GB table). Currently ~24% filled.
OPcache 256 MB: KEEP (runtime used unknown; config is not tight on paper).
SSR: ~116 MB RSS; not a memory problem. Frequent SIGTERM restarts are operator/supervisor, not OOM.
Redis: 879 MB used / 2 GB max; RSS 167 MB — likely swapped. 0 evictions.
Largest memory consumers are MySQL (8.1 GB RSS), vision Docker/Python (~2.5 GB), ClamAV (1.1 GB) — not Skinbase FPM.
1. Server RAM / swap
| Metric | Value |
|---|---|
| MemTotal | 23 GiB (24,615,736 kB) |
| Mem used | 15 GiB |
| MemAvailable | 8.2 GiB |
| Buffers+Cached | 7.2 GiB |
| AnonPages | 13.4 GiB |
| SwapTotal | 2.0 GiB |
| Swap used | 2.0 GiB (37 MiB free) |
| SwapCached | 774 MiB |
| Committed_AS | 30.1 GiB |
| vm.swappiness | 10 |
| vm.overcommit_memory | 0 |
| Load | 5.09 / 7.07 / 8.33 (8 CPUs) |
| PSI memory | some/full avg10=0.00 |
| PSI cpu | some avg10=4.7, avg60=14.5 |
| vmstat si/so (1s×5) | 0 / 0 |
| Processes | 243 |
Swap interpretation: cold / previously used pages sitting in swap, with 774 MiB SwapCached. Not active thrashing. Do not disable swap. Do not resize without a RAM plan. swappiness=10 is already conservative.
2. Process memory (RSS, measured)
| Process | RSS | Notes |
|---|---|---|
| mysqld | 8087 MB | InnoDB pool configured 10 GB |
| uvicorn/python (vision, port 8000) | 1310 + 542 + 238 + 197 + 103 + 54 MB | Docker vision stack |
| clamd | 1107 MB | |
| qdrant | 262 MB | vision |
| meilisearch | 395 MB | Skinbase search |
| crowdsec | 296 MB | |
| redis-server | 171 MB RSS / 879 MB used_memory |
swapped |
| netdata | 180 MB | |
| gitea | 156 MB | |
| php-fpm master | 48 MB | |
| php-fpm pool skinbase (×6) | 88–110 MB each | see §4 |
| php-fpm pool www (×2) | 13 MB each | unused by Skinbase vhost |
| horizon master | 104 MB | |
| horizon supervisors (×3) | ~103 MB each | |
| horizon workers (×6 now) | 103–125 MB | --memory=128 |
| node SSR | 116 MB | uptime minutes (restarted) |
| reverb | 106 MB | |
| nginx workers | 10–58 MB | one “shutting down” 15d |
| php8.4-fpm cgroup | current 746 MB, peak 1575 MB | all pools |
3. PHP-FPM Skinbase pool (php8.4-fpm-skinbase.sock)
From /etc/php/8.4/fpm/pool.d/skinbase.conf:
| Setting | Production value |
|---|---|
| pm | dynamic |
| pm.max_children | 14 |
| pm.start_servers | 4 |
| pm.min_spare_servers | 2 |
| pm.max_spare_servers | 6 |
| pm.process_idle_timeout | unset (N/A for dynamic) |
| pm.max_requests | 500 |
| request_terminate_timeout | 120s |
| request_slowlog_timeout | 10s |
| slowlog | /var/log/php8.4-fpm-skinbase-slow.log |
| pm.status_path | /fpm-status (localhost) |
| ping.path | /fpm-ping |
| listen.backlog | unset → PHP default 511 |
| memory_limit | 256M (pool admin) |
| max_execution_time | 60 |
Separate www pool: pm.max_children=5 on /run/php/php8.4-fpm.sock.
4. FPM status (localhost /fpm-status, pool up ~28.7h since 2026-08-22 10:50)
| Field | Value |
|---|---|
| accepted conn | 214,108 |
| listen queue | 0 |
| max listen queue | 0 |
| listen queue len | 0 |
| idle processes | 5 |
| active processes | 1 |
| total processes | 6 |
| max active processes | 14 |
| max children reached | 10 |
| slow requests (since start) | 2 |
| memory peak (FPM counter) | 60 MB |
Not currently exhausting workers. Peak has used all 14 children. No socket backlog was recorded, so nginx did not sit behind a full listen queue — capacity is tight at peak, idle most of the time.
Worker RSS (KB): 87712, 89996, 106444, 109308, 109476, 109480.
| Stat | RSS |
|---|---|
| min | 86 MB |
| median | 108 MB |
| mean | 102 MB |
| max | 110 MB |
| PSS | not readable (permission) |
| oldest worker | ~7.5 min (max_requests=500 recycling) |
memory_limit 256M is the ceiling, not typical RSS. Do not size max_children from 256M × 14.
Safe children from measured RSS:
typical_rss ≈ 110 MB
14 × 110 MB ≈ 1.54 GB (matches systemd MemoryPeak 1.58 GB for all FPM)
14 × 256 MB ≈ 3.58 GB worst-case if every worker hits the PHP limit
Raising max_children on a host with swap already full and AnonPages 13.4 GiB is not justified. Keep 14.
5. OPcache (config; FPM runtime stats not exposed)
| Setting | Value |
|---|---|
| enable | 1 |
| memory_consumption | 256 MB |
| interned_strings_buffer | 32 MB |
| max_accelerated_files | 50,000 |
| validate_timestamps | 1 |
| revalidate_freq | 2 s |
| JIT | disable / 0 |
CLI cannot read the FPM cache. 50k file slots and 256 MB are large for this Laravel app. Do not increase. Optional later (MEDIUM, not memory): validate_timestamps=0 after atomic releases.
6. Horizon (Supervisor skinbase-horizon)
Production env (config/horizon.php on the server):
| Supervisor | queues | maxProcesses | timeout | tries | memory |
|---|---|---|---|---|---|
| supervisor-default | search, default | 5 | 960 | 1 | 128 |
| supervisor-messaging | broadcasts, notifications | 3 | 90 | 1 | 128 |
| supervisor-mail | 2 | 90 | 5 | 128 |
Measured now: 1 master + 3 supervisors + 6 workers ≈ 1.15 GB RSS.
Peak if all maxProcesses spawn: 1+3+5+3+2 = 14 PHP procs × ~110 MB ≈ 1.5 GB.
Mail isolation is intact (tries=5, own supervisor). Do not merge queues.
Redis lists (prefix as used by Laravel):
| List | LLEN |
|---|---|
| queues:default | 796 |
| queues:mail | 0 |
| queues:search | 0 |
| queues:broadcasts | 0 |
| queues:notifications | 0 |
HIGH: default queue depth 796 — worker throughput, not FPM RAM. Investigate job mix in a later milestone; do not add Horizon processes until RAM is free (each extra worker ≈ 110 MB and --memory=128 is already near RSS 125 MB on default).
Horizon last supervisor restart: 2026-08-23 14:36 (clean exit 0 / SIGTERM wait), not OOM.
7. SSR
- RSS 116 MB, Node
/opt/www/virtual/SkinbaseNova/bootstrap/ssr/ssr.js - Supervisor restarts today: 14:22, 14:30, 14:47, 15:14, 15:16, 15:28 — all SIGTERM “waiting to stop”, then spawn. Not crash loops from memory.
- Heap not sampled (would require attaching to Node).
- Not material to host memory pressure.
8. Redis (read-only via Laravel)
| Field | Value |
|---|---|
| used_memory | 879 MB (peak 917 MB) |
| used_memory_rss | 167 MB |
| maxmemory | 2.00 GB |
| policy | allkeys-lru |
| fragmentation_ratio | 0.19 |
| evicted_keys | 0 |
| expired_keys | 225,783 |
| connected_clients | 28 |
| blocked_clients | 0 |
| keys | 3773 (3026 with TTL) |
| AOF | off |
| ops/sec | ~50 |
| hit/miss | 1.46M / 1.39M |
frag 0.19 + RSS << used_memory ⇒ Redis pages are in swap. Latency risk under cache bursts. 0 evictions: cache is within 2 GB. Do not flush. Do not raise maxmemory.
9. MySQL / Percona (read-only)
| Field | Value |
|---|---|
| innodb_buffer_pool_size | 10.00 GB (10 instances) |
| pages_data / total | 156,554 / 655,360 = 23.9% |
| pages_free | 498,726 |
| pool hit | 1 − 104,104 / 23,840,832,084 ≈ 99.9996% |
| max_connections | 120 |
| Max_used_connections | 26 |
| Threads_connected / running | 12 / 3 |
| tmp tables / disk tmp | 1,987,237 / 21 |
| innodb_redo_log_capacity | 2 GB |
| Slow_queries | 11,997 (uptime 103,663 s) |
10 GB pool is oversized for today’s ~2.8 GB schema, but M5 30-day hourly snapshots grow toward ~7.5 GB. Keeping 10 GB avoids refitting later. The table does not need to sit entirely in the pool; hit ratio is already excellent. Do not shrink now; do not grow.
10. Combined memory budget (measured / plausible peak)
| Bucket | Steady now | Plausible peak |
|---|---|---|
| OS/page cache (reclaimable) | ~7 GB cache | shrinks under pressure |
| MySQL | 8.1 GB | ~10 GB pool |
| Vision Docker/Python/Qdrant | ~2.7 GB | similar |
| ClamAV | 1.1 GB | similar |
| Meilisearch | 0.40 GB | 0.5 GB |
| Crowdsec | 0.30 GB | similar |
| Redis RSS | 0.17 GB | 2 GB if unswapped |
| PHP-FPM Skinbase | 0.65 GB (6×110) | 1.54 GB (14×110) |
| PHP-FPM www | 0.03 GB | 0.06 GB |
| Horizon | 1.15 GB | 1.5 GB |
| SSR + Reverb | 0.22 GB | 0.3 GB |
| nginx | 0.20 GB | 0.25 GB |
| Netdata/Gitea/journald/docker | ~0.5 GB | similar |
| Swap | 2 GB in use | — |
Steady anonymous ~14 GB + 2 GB swap explains the full swap file while 8.2 GB still Available as cache. Peak Skinbase (FPM 14 + Horizon 14) ≈ 3.0 GB, already observed in FPM cgroup peak 1.58 GB.
11. PHP-FPM recommendation
NO CHANGE to pool sizing.
pm = dynamic
pm.max_children = 14 # keep; peak already 14, listen queue 0
pm.start_servers = 4 # keep
pm.min_spare_servers = 2 # keep
pm.max_spare_servers = 6 # keep (matches current total=6)
pm.max_requests = 500 # keep; workers recycle in minutes
Calculation: median RSS 108 MB; 14×108≈1.5 GB. Available RAM is cache, not idle anonymous. Raising children increases swap of Redis/MySQL. Lowering below 14 would clip the measured peak (max active processes=14).
12. Swap recommendation
| Option | Decision |
|---|---|
| Keep 2 GiB swap | YES |
| Resize swap | NO (overflow is working; host needs RAM or fewer co-tenants) |
| Change vm.swappiness | NO (already 10) |
| Disable swap | NO |
13. OOM / crash safety
| Check | Result |
|---|---|
Kernel OOM (journalctl -k / dmesg) |
no lines visible to this user |
| PHP-FPM worker memory | YES — Allowed memory size of 268435456 Predis StreamConnection.php 2026-08-13 and 2026-08-18 |
| Redis OOM / evictions | no (evicted_keys=0) |
| MySQL | no evidence in readable logs |
| Horizon | clean supervisor stop 14:36, not killed |
| SSR | SIGTERM restarts, not OOM |
14. Ranked findings
| Sev | Finding | Change now? |
|---|---|---|
| HIGH | Swap 2 GiB ~full; Redis RSS≪used_memory | No sysctl. Free RAM (vision/clam) or larger VM later |
| HIGH | queues:default LLEN=796 |
Not FPM. Later queue milestone; don’t add Horizon RAM |
| HIGH | Shared host: vision ~2.7 GB + clamd 1.1 GB | Out of Skinbase app; ops placement |
| MEDIUM | FPM max_children reached 10×, max active 14 |
Keep 14; watch listen queue |
| MEDIUM | PHP 256 MB fatals on Predis | App/job payload size, not pool size |
| MEDIUM | slowlog 10s, 139k historical lines; status slow=2 since 22 Aug | Leave timeout; log rotation ops |
| LOW | nginx worker “shutting down” 15d | Recycle nginx in a maintenance window |
| LOW | www pool unused by Skinbase vhost | Optional disable later |
| NO CHANGE | OPcache 256/32/50k | |
| NO CHANGE | InnoDB 10 GB | Needed as snapshots grow |
| NO CHANGE | Horizon isolation / mail supervisor | |
| NO CHANGE | SSR RSS |
15. Expected benefit of doing nothing to FPM
Avoid swapping MySQL/Redis further. Current request path is not FPM-bound (active=1, listen queue=0).
16. Deployment plan (only if a change is later approved)
M6 recommends no FPM/swap/sysctl/MySQL/Horizon count change.
If ops later moves vision off-box or adds RAM:
- Re-measure RSS and
/fpm-status. - Then consider
pm.max_childrenusing new median RSS, not 256M. - Pool edit +
systemctl reload php8.4-fpm(reload, not restart if possible). - Rollback: restore
skinbase.confand reload.
Any pm.* or memory_limit change requires FPM reload. Swap/sysctl/MySQL buffer pool need service-level restarts — not proposed.
17. Risks of increasing max_children without more RAM
More anonymous PHP, more swap of Redis (cache latency) and InnoDB (buffer pool eviction), possible FPM 256 MB fatals still happen per-request.
Production was not modified.