Files
SkinbaseNova/docs/optimization-m6-php-fpm-memory.md
T
klevze 8a80aae21e Ship production optimization M1-M12.5A: queues, metrics, HTTP observability, and vector search reliability.
Keep similar-ai from tripping the global circuit on a lone URL 502, clamp Qdrant search to 100, and add Server-Timing plus slow-request logging. Studio shared props, Academy S3 exists caching, heat chunking, and Redis/scheduler hygiene stay in this rollout.
2026-08-25 07:58:47 +02:00

358 lines
12 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# M6 — PHP-FPM, Memory, Swap & Worker Capacity
```text
STATUS: COMPLETE (read-only; no production changes)
WHEN: 2026-08-23 15:29–15:32 Europe/Ljubljana (UTC 13:29)
HOST: server3, Debian 13, kernel 6.12.96, 8 cores, uptime 24d 18h
```
Nothing was restarted, tuned, flushed, or edited on server3.
---
## Verdict
```text
Swap: KEEP 2 GiB. Nearly full, but NOT thrashing (si/so=0, PSI memory=0).
PHP-FPM max_children=14: KEEP. Peak has used all 14; listen queue never built.
Do not raise FPM/Horizon on this host until RAM is freed or the machine is larger.
InnoDB 10 GB pool: KEEP for 30-day snapshot growth (~7.5 GB table). Currently ~24% filled.
OPcache 256 MB: KEEP (runtime used unknown; config is not tight on paper).
SSR: ~116 MB RSS; not a memory problem. Frequent SIGTERM restarts are operator/supervisor, not OOM.
Redis: 879 MB used / 2 GB max; RSS 167 MB — likely swapped. 0 evictions.
```
Largest memory consumers are **MySQL (8.1 GB RSS)**, **vision Docker/Python (~2.5 GB)**, **ClamAV (1.1 GB)** — not Skinbase FPM.
---
## 1. Server RAM / swap
| Metric | Value |
| ------ | ----- |
| MemTotal | 23 GiB (24,615,736 kB) |
| Mem used | 15 GiB |
| MemAvailable | **8.2 GiB** |
| Buffers+Cached | 7.2 GiB |
| AnonPages | 13.4 GiB |
| SwapTotal | 2.0 GiB |
| Swap used | **2.0 GiB (37 MiB free)** |
| SwapCached | 774 MiB |
| Committed_AS | 30.1 GiB |
| vm.swappiness | 10 |
| vm.overcommit_memory | 0 |
| Load | 5.09 / 7.07 / 8.33 (8 CPUs) |
| PSI memory | some/full avg10=**0.00** |
| PSI cpu | some avg10=4.7, avg60=14.5 |
| vmstat si/so (1s×5) | **0 / 0** |
| Processes | 243 |
**Swap interpretation:** cold / previously used pages sitting in swap, with **774 MiB SwapCached**. Not active thrashing. Do **not** disable swap. Do **not** resize without a RAM plan. swappiness=10 is already conservative.
---
## 2. Process memory (RSS, measured)
| Process | RSS | Notes |
| ------- | --: | ----- |
| mysqld | **8087 MB** | InnoDB pool configured 10 GB |
| uvicorn/python (vision, port 8000) | **1310 + 542 + 238 + 197 + 103 + 54 MB** | Docker vision stack |
| clamd | **1107 MB** | |
| qdrant | 262 MB | vision |
| meilisearch | 395 MB | Skinbase search |
| crowdsec | 296 MB | |
| redis-server | 171 MB RSS / **879 MB** `used_memory` | swapped |
| netdata | 180 MB | |
| gitea | 156 MB | |
| php-fpm master | 48 MB | |
| php-fpm pool skinbase (×6) | **88–110 MB** each | see §4 |
| php-fpm pool www (×2) | 13 MB each | unused by Skinbase vhost |
| horizon master | 104 MB | |
| horizon supervisors (×3) | ~103 MB each | |
| horizon workers (×6 now) | 103–125 MB | `--memory=128` |
| node SSR | 116 MB | uptime minutes (restarted) |
| reverb | 106 MB | |
| nginx workers | 10–58 MB | one “shutting down” 15d |
| php8.4-fpm cgroup | current **746 MB**, peak **1575 MB** | all pools |
---
## 3. PHP-FPM Skinbase pool (`php8.4-fpm-skinbase.sock`)
From `/etc/php/8.4/fpm/pool.d/skinbase.conf`:
| Setting | Production value |
| ------- | ---------------- |
| pm | **dynamic** |
| pm.max_children | **14** |
| pm.start_servers | 4 |
| pm.min_spare_servers | 2 |
| pm.max_spare_servers | 6 |
| pm.process_idle_timeout | unset (N/A for dynamic) |
| pm.max_requests | **500** |
| request_terminate_timeout | 120s |
| request_slowlog_timeout | 10s |
| slowlog | `/var/log/php8.4-fpm-skinbase-slow.log` |
| pm.status_path | `/fpm-status` (localhost) |
| ping.path | `/fpm-ping` |
| listen.backlog | unset → PHP default **511** |
| memory_limit | **256M** (pool admin) |
| max_execution_time | 60 |
Separate `www` pool: `pm.max_children=5` on `/run/php/php8.4-fpm.sock`.
---
## 4. FPM status (localhost `/fpm-status`, pool up ~28.7h since 2026-08-22 10:50)
| Field | Value |
| ----- | ----- |
| accepted conn | 214,108 |
| listen queue | **0** |
| max listen queue | **0** |
| listen queue len | 0 |
| idle processes | 5 |
| active processes | 1 |
| total processes | 6 |
| max active processes | **14** |
| max children reached | **10** |
| slow requests (since start) | **2** |
| memory peak (FPM counter) | 60 MB |
**Not currently exhausting workers.** Peak has used all 14 children. No socket backlog was recorded, so nginx did not sit behind a full listen queue — capacity is **tight at peak, idle most of the time**.
Worker RSS (KB): 87712, 89996, 106444, 109308, 109476, 109480.
| Stat | RSS |
| ---- | --: |
| min | 86 MB |
| median | **108 MB** |
| mean | **102 MB** |
| max | **110 MB** |
| PSS | not readable (permission) |
| oldest worker | ~7.5 min (max_requests=500 recycling) |
`memory_limit` 256M is the ceiling, not typical RSS. **Do not size max_children from 256M × 14.**
Safe children from measured RSS:
```text
typical_rss ≈ 110 MB
14 × 110 MB ≈ 1.54 GB (matches systemd MemoryPeak 1.58 GB for all FPM)
14 × 256 MB ≈ 3.58 GB worst-case if every worker hits the PHP limit
```
Raising max_children on a host with **swap already full** and AnonPages 13.4 GiB is not justified. **Keep 14.**
---
## 5. OPcache (config; FPM runtime stats not exposed)
| Setting | Value |
| ------- | ----- |
| enable | 1 |
| memory_consumption | 256 MB |
| interned_strings_buffer | 32 MB |
| max_accelerated_files | 50,000 |
| validate_timestamps | 1 |
| revalidate_freq | 2 s |
| JIT | disable / 0 |
CLI cannot read the FPM cache. 50k file slots and 256 MB are large for this Laravel app. **Do not increase.** Optional later (MEDIUM, not memory): `validate_timestamps=0` after atomic releases.
---
## 6. Horizon (Supervisor `skinbase-horizon`)
Production env (`config/horizon.php` on the server):
| Supervisor | queues | maxProcesses | timeout | tries | memory |
| ---------- | ------ | ------------: | ------: | ----: | -----: |
| supervisor-default | search, default | 5 | 960 | 1 | 128 |
| supervisor-messaging | broadcasts, notifications | 3 | 90 | 1 | 128 |
| supervisor-mail | mail | 2 | 90 | 5 | 128 |
Measured now: 1 master + 3 supervisors + 6 workers ≈ **1.15 GB RSS**.
Peak if all maxProcesses spawn: 1+3+5+3+2 = **14 PHP procs × ~110 MB ≈ 1.5 GB**.
Mail isolation is intact (`tries=5`, own supervisor). **Do not merge queues.**
Redis lists (prefix as used by Laravel):
| List | LLEN |
| ---- | ---: |
| queues:default | **796** |
| queues:mail | 0 |
| queues:search | 0 |
| queues:broadcasts | 0 |
| queues:notifications | 0 |
**HIGH:** default queue depth 796 — worker **throughput**, not FPM RAM. Investigate job mix in a later milestone; do not add Horizon processes until RAM is free (each extra worker ≈ 110 MB and `--memory=128` is already near RSS 125 MB on default).
Horizon last supervisor restart: 2026-08-23 14:36 (clean exit 0 / SIGTERM wait), not OOM.
---
## 7. SSR
- RSS **116 MB**, Node `/opt/www/virtual/SkinbaseNova/bootstrap/ssr/ssr.js`
- Supervisor restarts today: 14:22, 14:30, 14:47, 15:14, 15:16, 15:28 — all **SIGTERM** “waiting to stop”, then spawn. Not crash loops from memory.
- Heap not sampled (would require attaching to Node).
- **Not material** to host memory pressure.
---
## 8. Redis (read-only via Laravel)
| Field | Value |
| ----- | ----- |
| used_memory | 879 MB (peak 917 MB) |
| used_memory_rss | **167 MB** |
| maxmemory | 2.00 GB |
| policy | allkeys-lru |
| fragmentation_ratio | **0.19** |
| evicted_keys | **0** |
| expired_keys | 225,783 |
| connected_clients | 28 |
| blocked_clients | 0 |
| keys | 3773 (3026 with TTL) |
| AOF | off |
| ops/sec | ~50 |
| hit/miss | 1.46M / 1.39M |
frag 0.19 + RSS << used_memory ⇒ **Redis pages are in swap**. Latency risk under cache bursts. 0 evictions: cache is within 2 GB. Do not flush. Do not raise maxmemory.
---
## 9. MySQL / Percona (read-only)
| Field | Value |
| ----- | ----- |
| innodb_buffer_pool_size | **10.00 GB** (10 instances) |
| pages_data / total | 156,554 / 655,360 = **23.9%** |
| pages_free | 498,726 |
| pool hit | 1 − 104,104 / 23,840,832,084 ≈ **99.9996%** |
| max_connections | 120 |
| Max_used_connections | 26 |
| Threads_connected / running | 12 / 3 |
| tmp tables / disk tmp | 1,987,237 / **21** |
| innodb_redo_log_capacity | 2 GB |
| Slow_queries | 11,997 (uptime 103,663 s) |
10 GB pool is **oversized for today’s ~2.8 GB schema**, but M5 30-day hourly snapshots grow toward **~7.5 GB**. Keeping 10 GB avoids refitting later. The table does **not** need to sit entirely in the pool; hit ratio is already excellent. **Do not shrink now; do not grow.**
---
## 10. Combined memory budget (measured / plausible peak)
| Bucket | Steady now | Plausible peak |
| ------ | ---------: | -------------: |
| OS/page cache (reclaimable) | ~7 GB cache | shrinks under pressure |
| MySQL | 8.1 GB | ~10 GB pool |
| Vision Docker/Python/Qdrant | ~2.7 GB | similar |
| ClamAV | 1.1 GB | similar |
| Meilisearch | 0.40 GB | 0.5 GB |
| Crowdsec | 0.30 GB | similar |
| Redis RSS | 0.17 GB | 2 GB if unswapped |
| PHP-FPM Skinbase | 0.65 GB (6×110) | **1.54 GB** (14×110) |
| PHP-FPM www | 0.03 GB | 0.06 GB |
| Horizon | 1.15 GB | **1.5 GB** |
| SSR + Reverb | 0.22 GB | 0.3 GB |
| nginx | 0.20 GB | 0.25 GB |
| Netdata/Gitea/journald/docker | ~0.5 GB | similar |
| Swap | 2 GB **in use** | — |
Steady anonymous ~14 GB + 2 GB swap explains the full swap file while **8.2 GB still Available** as cache. Peak Skinbase (FPM 14 + Horizon 14) ≈ **3.0 GB**, already observed in FPM cgroup peak 1.58 GB.
---
## 11. PHP-FPM recommendation
**NO CHANGE** to pool sizing.
```text
pm = dynamic
pm.max_children = 14 # keep; peak already 14, listen queue 0
pm.start_servers = 4 # keep
pm.min_spare_servers = 2 # keep
pm.max_spare_servers = 6 # keep (matches current total=6)
pm.max_requests = 500 # keep; workers recycle in minutes
```
Calculation: median RSS 108 MB; 14×108≈1.5 GB. Available RAM is cache, not idle anonymous. Raising children increases swap of Redis/MySQL. Lowering below 14 would clip the measured peak (`max active processes=14`).
---
## 12. Swap recommendation
| Option | Decision |
| ------ | -------- |
| Keep 2 GiB swap | **YES** |
| Resize swap | **NO** (overflow is working; host needs RAM or fewer co-tenants) |
| Change vm.swappiness | **NO** (already 10) |
| Disable swap | **NO** |
---
## 13. OOM / crash safety
| Check | Result |
| ----- | ------ |
| Kernel OOM (`journalctl -k` / dmesg) | **no lines** visible to this user |
| PHP-FPM worker memory | **YES** — `Allowed memory size of 268435456` Predis `StreamConnection.php` 2026-08-13 and 2026-08-18 |
| Redis OOM / evictions | **no** (evicted_keys=0) |
| MySQL | no evidence in readable logs |
| Horizon | clean supervisor stop 14:36, not killed |
| SSR | SIGTERM restarts, not OOM |
---
## 14. Ranked findings
| Sev | Finding | Change now? |
| --- | ------- | ----------- |
| HIGH | Swap 2 GiB ~full; Redis RSS≪used_memory | No sysctl. Free RAM (vision/clam) or larger VM later |
| HIGH | `queues:default` LLEN=**796** | Not FPM. Later queue milestone; don’t add Horizon RAM |
| HIGH | Shared host: vision ~2.7 GB + clamd 1.1 GB | Out of Skinbase app; ops placement |
| MEDIUM | FPM `max_children` reached 10×, max active 14 | Keep 14; watch listen queue |
| MEDIUM | PHP 256 MB fatals on Predis | App/job payload size, not pool size |
| MEDIUM | slowlog 10s, 139k historical lines; status slow=2 since 22 Aug | Leave timeout; log rotation ops |
| LOW | nginx worker “shutting down” 15d | Recycle nginx in a maintenance window |
| LOW | www pool unused by Skinbase vhost | Optional disable later |
| NO CHANGE | OPcache 256/32/50k | |
| NO CHANGE | InnoDB 10 GB | Needed as snapshots grow |
| NO CHANGE | Horizon isolation / mail supervisor | |
| NO CHANGE | SSR RSS | |
---
## 15. Expected benefit of doing nothing to FPM
Avoid swapping MySQL/Redis further. Current request path is not FPM-bound (`active=1`, `listen queue=0`).
---
## 16. Deployment plan (only if a change is later approved)
M6 recommends **no FPM/swap/sysctl/MySQL/Horizon count change**.
If ops later moves vision off-box or adds RAM:
1. Re-measure RSS and `/fpm-status`.
2. Then consider `pm.max_children` using **new** median RSS, not 256M.
3. Pool edit + `systemctl reload php8.4-fpm` (reload, not restart if possible).
4. Rollback: restore `skinbase.conf` and reload.
Any `pm.*` or `memory_limit` change **requires FPM reload**. Swap/sysctl/MySQL buffer pool need service-level restarts — not proposed.
---
## 17. Risks of increasing max_children without more RAM
More anonymous PHP, more swap of Redis (cache latency) and InnoDB (buffer pool eviction), possible FPM 256 MB fatals still happen per-request.
---
Production was not modified.