Keep similar-ai from tripping the global circuit on a lone URL 502, clamp Qdrant search to 100, and add Server-Timing plus slow-request logging. Studio shared props, Academy S3 exists caching, heat chunking, and Redis/scheduler hygiene stay in this rollout.
358 lines
12 KiB
Markdown
358 lines
12 KiB
Markdown
# M6 — PHP-FPM, Memory, Swap & Worker Capacity
|
||
|
||
```text
|
||
STATUS: COMPLETE (read-only; no production changes)
|
||
WHEN: 2026-08-23 15:29–15:32 Europe/Ljubljana (UTC 13:29)
|
||
HOST: server3, Debian 13, kernel 6.12.96, 8 cores, uptime 24d 18h
|
||
```
|
||
|
||
Nothing was restarted, tuned, flushed, or edited on server3.
|
||
|
||
---
|
||
|
||
## Verdict
|
||
|
||
```text
|
||
Swap: KEEP 2 GiB. Nearly full, but NOT thrashing (si/so=0, PSI memory=0).
|
||
PHP-FPM max_children=14: KEEP. Peak has used all 14; listen queue never built.
|
||
Do not raise FPM/Horizon on this host until RAM is freed or the machine is larger.
|
||
InnoDB 10 GB pool: KEEP for 30-day snapshot growth (~7.5 GB table). Currently ~24% filled.
|
||
OPcache 256 MB: KEEP (runtime used unknown; config is not tight on paper).
|
||
SSR: ~116 MB RSS; not a memory problem. Frequent SIGTERM restarts are operator/supervisor, not OOM.
|
||
Redis: 879 MB used / 2 GB max; RSS 167 MB — likely swapped. 0 evictions.
|
||
```
|
||
|
||
Largest memory consumers are **MySQL (8.1 GB RSS)**, **vision Docker/Python (~2.5 GB)**, **ClamAV (1.1 GB)** — not Skinbase FPM.
|
||
|
||
---
|
||
|
||
## 1. Server RAM / swap
|
||
|
||
| Metric | Value |
|
||
| ------ | ----- |
|
||
| MemTotal | 23 GiB (24,615,736 kB) |
|
||
| Mem used | 15 GiB |
|
||
| MemAvailable | **8.2 GiB** |
|
||
| Buffers+Cached | 7.2 GiB |
|
||
| AnonPages | 13.4 GiB |
|
||
| SwapTotal | 2.0 GiB |
|
||
| Swap used | **2.0 GiB (37 MiB free)** |
|
||
| SwapCached | 774 MiB |
|
||
| Committed_AS | 30.1 GiB |
|
||
| vm.swappiness | 10 |
|
||
| vm.overcommit_memory | 0 |
|
||
| Load | 5.09 / 7.07 / 8.33 (8 CPUs) |
|
||
| PSI memory | some/full avg10=**0.00** |
|
||
| PSI cpu | some avg10=4.7, avg60=14.5 |
|
||
| vmstat si/so (1s×5) | **0 / 0** |
|
||
| Processes | 243 |
|
||
|
||
**Swap interpretation:** cold / previously used pages sitting in swap, with **774 MiB SwapCached**. Not active thrashing. Do **not** disable swap. Do **not** resize without a RAM plan. swappiness=10 is already conservative.
|
||
|
||
---
|
||
|
||
## 2. Process memory (RSS, measured)
|
||
|
||
| Process | RSS | Notes |
|
||
| ------- | --: | ----- |
|
||
| mysqld | **8087 MB** | InnoDB pool configured 10 GB |
|
||
| uvicorn/python (vision, port 8000) | **1310 + 542 + 238 + 197 + 103 + 54 MB** | Docker vision stack |
|
||
| clamd | **1107 MB** | |
|
||
| qdrant | 262 MB | vision |
|
||
| meilisearch | 395 MB | Skinbase search |
|
||
| crowdsec | 296 MB | |
|
||
| redis-server | 171 MB RSS / **879 MB** `used_memory` | swapped |
|
||
| netdata | 180 MB | |
|
||
| gitea | 156 MB | |
|
||
| php-fpm master | 48 MB | |
|
||
| php-fpm pool skinbase (×6) | **88–110 MB** each | see §4 |
|
||
| php-fpm pool www (×2) | 13 MB each | unused by Skinbase vhost |
|
||
| horizon master | 104 MB | |
|
||
| horizon supervisors (×3) | ~103 MB each | |
|
||
| horizon workers (×6 now) | 103–125 MB | `--memory=128` |
|
||
| node SSR | 116 MB | uptime minutes (restarted) |
|
||
| reverb | 106 MB | |
|
||
| nginx workers | 10–58 MB | one “shutting down” 15d |
|
||
| php8.4-fpm cgroup | current **746 MB**, peak **1575 MB** | all pools |
|
||
|
||
---
|
||
|
||
## 3. PHP-FPM Skinbase pool (`php8.4-fpm-skinbase.sock`)
|
||
|
||
From `/etc/php/8.4/fpm/pool.d/skinbase.conf`:
|
||
|
||
| Setting | Production value |
|
||
| ------- | ---------------- |
|
||
| pm | **dynamic** |
|
||
| pm.max_children | **14** |
|
||
| pm.start_servers | 4 |
|
||
| pm.min_spare_servers | 2 |
|
||
| pm.max_spare_servers | 6 |
|
||
| pm.process_idle_timeout | unset (N/A for dynamic) |
|
||
| pm.max_requests | **500** |
|
||
| request_terminate_timeout | 120s |
|
||
| request_slowlog_timeout | 10s |
|
||
| slowlog | `/var/log/php8.4-fpm-skinbase-slow.log` |
|
||
| pm.status_path | `/fpm-status` (localhost) |
|
||
| ping.path | `/fpm-ping` |
|
||
| listen.backlog | unset → PHP default **511** |
|
||
| memory_limit | **256M** (pool admin) |
|
||
| max_execution_time | 60 |
|
||
|
||
Separate `www` pool: `pm.max_children=5` on `/run/php/php8.4-fpm.sock`.
|
||
|
||
---
|
||
|
||
## 4. FPM status (localhost `/fpm-status`, pool up ~28.7h since 2026-08-22 10:50)
|
||
|
||
| Field | Value |
|
||
| ----- | ----- |
|
||
| accepted conn | 214,108 |
|
||
| listen queue | **0** |
|
||
| max listen queue | **0** |
|
||
| listen queue len | 0 |
|
||
| idle processes | 5 |
|
||
| active processes | 1 |
|
||
| total processes | 6 |
|
||
| max active processes | **14** |
|
||
| max children reached | **10** |
|
||
| slow requests (since start) | **2** |
|
||
| memory peak (FPM counter) | 60 MB |
|
||
|
||
**Not currently exhausting workers.** Peak has used all 14 children. No socket backlog was recorded, so nginx did not sit behind a full listen queue — capacity is **tight at peak, idle most of the time**.
|
||
|
||
Worker RSS (KB): 87712, 89996, 106444, 109308, 109476, 109480.
|
||
|
||
| Stat | RSS |
|
||
| ---- | --: |
|
||
| min | 86 MB |
|
||
| median | **108 MB** |
|
||
| mean | **102 MB** |
|
||
| max | **110 MB** |
|
||
| PSS | not readable (permission) |
|
||
| oldest worker | ~7.5 min (max_requests=500 recycling) |
|
||
|
||
`memory_limit` 256M is the ceiling, not typical RSS. **Do not size max_children from 256M × 14.**
|
||
|
||
Safe children from measured RSS:
|
||
|
||
```text
|
||
typical_rss ≈ 110 MB
|
||
14 × 110 MB ≈ 1.54 GB (matches systemd MemoryPeak 1.58 GB for all FPM)
|
||
14 × 256 MB ≈ 3.58 GB worst-case if every worker hits the PHP limit
|
||
```
|
||
|
||
Raising max_children on a host with **swap already full** and AnonPages 13.4 GiB is not justified. **Keep 14.**
|
||
|
||
---
|
||
|
||
## 5. OPcache (config; FPM runtime stats not exposed)
|
||
|
||
| Setting | Value |
|
||
| ------- | ----- |
|
||
| enable | 1 |
|
||
| memory_consumption | 256 MB |
|
||
| interned_strings_buffer | 32 MB |
|
||
| max_accelerated_files | 50,000 |
|
||
| validate_timestamps | 1 |
|
||
| revalidate_freq | 2 s |
|
||
| JIT | disable / 0 |
|
||
|
||
CLI cannot read the FPM cache. 50k file slots and 256 MB are large for this Laravel app. **Do not increase.** Optional later (MEDIUM, not memory): `validate_timestamps=0` after atomic releases.
|
||
|
||
---
|
||
|
||
## 6. Horizon (Supervisor `skinbase-horizon`)
|
||
|
||
Production env (`config/horizon.php` on the server):
|
||
|
||
| Supervisor | queues | maxProcesses | timeout | tries | memory |
|
||
| ---------- | ------ | ------------: | ------: | ----: | -----: |
|
||
| supervisor-default | search, default | 5 | 960 | 1 | 128 |
|
||
| supervisor-messaging | broadcasts, notifications | 3 | 90 | 1 | 128 |
|
||
| supervisor-mail | mail | 2 | 90 | 5 | 128 |
|
||
|
||
Measured now: 1 master + 3 supervisors + 6 workers ≈ **1.15 GB RSS**.
|
||
|
||
Peak if all maxProcesses spawn: 1+3+5+3+2 = **14 PHP procs × ~110 MB ≈ 1.5 GB**.
|
||
|
||
Mail isolation is intact (`tries=5`, own supervisor). **Do not merge queues.**
|
||
|
||
Redis lists (prefix as used by Laravel):
|
||
|
||
| List | LLEN |
|
||
| ---- | ---: |
|
||
| queues:default | **796** |
|
||
| queues:mail | 0 |
|
||
| queues:search | 0 |
|
||
| queues:broadcasts | 0 |
|
||
| queues:notifications | 0 |
|
||
|
||
**HIGH:** default queue depth 796 — worker **throughput**, not FPM RAM. Investigate job mix in a later milestone; do not add Horizon processes until RAM is free (each extra worker ≈ 110 MB and `--memory=128` is already near RSS 125 MB on default).
|
||
|
||
Horizon last supervisor restart: 2026-08-23 14:36 (clean exit 0 / SIGTERM wait), not OOM.
|
||
|
||
---
|
||
|
||
## 7. SSR
|
||
|
||
- RSS **116 MB**, Node `/opt/www/virtual/SkinbaseNova/bootstrap/ssr/ssr.js`
|
||
- Supervisor restarts today: 14:22, 14:30, 14:47, 15:14, 15:16, 15:28 — all **SIGTERM** “waiting to stop”, then spawn. Not crash loops from memory.
|
||
- Heap not sampled (would require attaching to Node).
|
||
- **Not material** to host memory pressure.
|
||
|
||
---
|
||
|
||
## 8. Redis (read-only via Laravel)
|
||
|
||
| Field | Value |
|
||
| ----- | ----- |
|
||
| used_memory | 879 MB (peak 917 MB) |
|
||
| used_memory_rss | **167 MB** |
|
||
| maxmemory | 2.00 GB |
|
||
| policy | allkeys-lru |
|
||
| fragmentation_ratio | **0.19** |
|
||
| evicted_keys | **0** |
|
||
| expired_keys | 225,783 |
|
||
| connected_clients | 28 |
|
||
| blocked_clients | 0 |
|
||
| keys | 3773 (3026 with TTL) |
|
||
| AOF | off |
|
||
| ops/sec | ~50 |
|
||
| hit/miss | 1.46M / 1.39M |
|
||
|
||
frag 0.19 + RSS << used_memory ⇒ **Redis pages are in swap**. Latency risk under cache bursts. 0 evictions: cache is within 2 GB. Do not flush. Do not raise maxmemory.
|
||
|
||
---
|
||
|
||
## 9. MySQL / Percona (read-only)
|
||
|
||
| Field | Value |
|
||
| ----- | ----- |
|
||
| innodb_buffer_pool_size | **10.00 GB** (10 instances) |
|
||
| pages_data / total | 156,554 / 655,360 = **23.9%** |
|
||
| pages_free | 498,726 |
|
||
| pool hit | 1 − 104,104 / 23,840,832,084 ≈ **99.9996%** |
|
||
| max_connections | 120 |
|
||
| Max_used_connections | 26 |
|
||
| Threads_connected / running | 12 / 3 |
|
||
| tmp tables / disk tmp | 1,987,237 / **21** |
|
||
| innodb_redo_log_capacity | 2 GB |
|
||
| Slow_queries | 11,997 (uptime 103,663 s) |
|
||
|
||
10 GB pool is **oversized for today’s ~2.8 GB schema**, but M5 30-day hourly snapshots grow toward **~7.5 GB**. Keeping 10 GB avoids refitting later. The table does **not** need to sit entirely in the pool; hit ratio is already excellent. **Do not shrink now; do not grow.**
|
||
|
||
---
|
||
|
||
## 10. Combined memory budget (measured / plausible peak)
|
||
|
||
| Bucket | Steady now | Plausible peak |
|
||
| ------ | ---------: | -------------: |
|
||
| OS/page cache (reclaimable) | ~7 GB cache | shrinks under pressure |
|
||
| MySQL | 8.1 GB | ~10 GB pool |
|
||
| Vision Docker/Python/Qdrant | ~2.7 GB | similar |
|
||
| ClamAV | 1.1 GB | similar |
|
||
| Meilisearch | 0.40 GB | 0.5 GB |
|
||
| Crowdsec | 0.30 GB | similar |
|
||
| Redis RSS | 0.17 GB | 2 GB if unswapped |
|
||
| PHP-FPM Skinbase | 0.65 GB (6×110) | **1.54 GB** (14×110) |
|
||
| PHP-FPM www | 0.03 GB | 0.06 GB |
|
||
| Horizon | 1.15 GB | **1.5 GB** |
|
||
| SSR + Reverb | 0.22 GB | 0.3 GB |
|
||
| nginx | 0.20 GB | 0.25 GB |
|
||
| Netdata/Gitea/journald/docker | ~0.5 GB | similar |
|
||
| Swap | 2 GB **in use** | — |
|
||
|
||
Steady anonymous ~14 GB + 2 GB swap explains the full swap file while **8.2 GB still Available** as cache. Peak Skinbase (FPM 14 + Horizon 14) ≈ **3.0 GB**, already observed in FPM cgroup peak 1.58 GB.
|
||
|
||
---
|
||
|
||
## 11. PHP-FPM recommendation
|
||
|
||
**NO CHANGE** to pool sizing.
|
||
|
||
```text
|
||
pm = dynamic
|
||
pm.max_children = 14 # keep; peak already 14, listen queue 0
|
||
pm.start_servers = 4 # keep
|
||
pm.min_spare_servers = 2 # keep
|
||
pm.max_spare_servers = 6 # keep (matches current total=6)
|
||
pm.max_requests = 500 # keep; workers recycle in minutes
|
||
```
|
||
|
||
Calculation: median RSS 108 MB; 14×108≈1.5 GB. Available RAM is cache, not idle anonymous. Raising children increases swap of Redis/MySQL. Lowering below 14 would clip the measured peak (`max active processes=14`).
|
||
|
||
---
|
||
|
||
## 12. Swap recommendation
|
||
|
||
| Option | Decision |
|
||
| ------ | -------- |
|
||
| Keep 2 GiB swap | **YES** |
|
||
| Resize swap | **NO** (overflow is working; host needs RAM or fewer co-tenants) |
|
||
| Change vm.swappiness | **NO** (already 10) |
|
||
| Disable swap | **NO** |
|
||
|
||
---
|
||
|
||
## 13. OOM / crash safety
|
||
|
||
| Check | Result |
|
||
| ----- | ------ |
|
||
| Kernel OOM (`journalctl -k` / dmesg) | **no lines** visible to this user |
|
||
| PHP-FPM worker memory | **YES** — `Allowed memory size of 268435456` Predis `StreamConnection.php` 2026-08-13 and 2026-08-18 |
|
||
| Redis OOM / evictions | **no** (evicted_keys=0) |
|
||
| MySQL | no evidence in readable logs |
|
||
| Horizon | clean supervisor stop 14:36, not killed |
|
||
| SSR | SIGTERM restarts, not OOM |
|
||
|
||
---
|
||
|
||
## 14. Ranked findings
|
||
|
||
| Sev | Finding | Change now? |
|
||
| --- | ------- | ----------- |
|
||
| HIGH | Swap 2 GiB ~full; Redis RSS≪used_memory | No sysctl. Free RAM (vision/clam) or larger VM later |
|
||
| HIGH | `queues:default` LLEN=**796** | Not FPM. Later queue milestone; don’t add Horizon RAM |
|
||
| HIGH | Shared host: vision ~2.7 GB + clamd 1.1 GB | Out of Skinbase app; ops placement |
|
||
| MEDIUM | FPM `max_children` reached 10×, max active 14 | Keep 14; watch listen queue |
|
||
| MEDIUM | PHP 256 MB fatals on Predis | App/job payload size, not pool size |
|
||
| MEDIUM | slowlog 10s, 139k historical lines; status slow=2 since 22 Aug | Leave timeout; log rotation ops |
|
||
| LOW | nginx worker “shutting down” 15d | Recycle nginx in a maintenance window |
|
||
| LOW | www pool unused by Skinbase vhost | Optional disable later |
|
||
| NO CHANGE | OPcache 256/32/50k | |
|
||
| NO CHANGE | InnoDB 10 GB | Needed as snapshots grow |
|
||
| NO CHANGE | Horizon isolation / mail supervisor | |
|
||
| NO CHANGE | SSR RSS | |
|
||
|
||
---
|
||
|
||
## 15. Expected benefit of doing nothing to FPM
|
||
|
||
Avoid swapping MySQL/Redis further. Current request path is not FPM-bound (`active=1`, `listen queue=0`).
|
||
|
||
---
|
||
|
||
## 16. Deployment plan (only if a change is later approved)
|
||
|
||
M6 recommends **no FPM/swap/sysctl/MySQL/Horizon count change**.
|
||
|
||
If ops later moves vision off-box or adds RAM:
|
||
|
||
1. Re-measure RSS and `/fpm-status`.
|
||
2. Then consider `pm.max_children` using **new** median RSS, not 256M.
|
||
3. Pool edit + `systemctl reload php8.4-fpm` (reload, not restart if possible).
|
||
4. Rollback: restore `skinbase.conf` and reload.
|
||
|
||
Any `pm.*` or `memory_limit` change **requires FPM reload**. Swap/sysctl/MySQL buffer pool need service-level restarts — not proposed.
|
||
|
||
---
|
||
|
||
## 17. Risks of increasing max_children without more RAM
|
||
|
||
More anonymous PHP, more swap of Redis (cache latency) and InnoDB (buffer pool eviction), possible FPM 256 MB fatals still happen per-request.
|
||
|
||
---
|
||
|
||
Production was not modified.
|