Ship production optimization M1-M12.5A: queues, metrics, HTTP observability, and vector search reliability.

Keep similar-ai from tripping the global circuit on a lone URL 502, clamp Qdrant search to 100, and add Server-Timing plus slow-request logging. Studio shared props, Academy S3 exists caching, heat chunking, and Redis/scheduler hygiene stay in this rollout.
This commit is contained in:
2026-08-25 07:58:47 +02:00
parent f52879edbb
commit 8a80aae21e
114 changed files with 9293 additions and 284 deletions
@@ -0,0 +1,186 @@
# M7 — Default Queue Backlog & Redis/PHP Memory Failures
```text
STATUS: COMPLETE (investigation + local smallest fixes; production unchanged)
WHEN: 2026-08-23 17:49 Europe/Ljubljana
```
No Horizon restart, Redis flush, job retry, memory_limit change, or worker-count change on server3.
---
## Verdict
```text
Default backlog class: C + D (producer > consumer; ranking/rec jobs block index jobs)
Default depth: 1208 → 1213 in 60s (net +5/min, growing)
Oldest default job: ~2.5 h
Search queue: 0 (idle workers exist on supervisor-default)
collections: 32,632 jobs, oldest 2026-04-30, NO Horizon consumer
forum-moderation: 552,888 AnalyzeForumPostJob, ~557 MB list, NO consumer
forum-security: 5,533 jobs, no consumer
presence index: SET 2,480,292 members, ~272 MB, TTL -1
Predis 256 MB fatals: loading those huge Redis replies over HTTP/FPM
Horizon workers: do NOT increase
search+default supervisor: KEEP together; put Index* jobs on search
```
---
## 1. `queues:default` contents (metadata only)
| Class | Count | Avg payload |
| ----- | ----: | ----------: |
| `RankBuildScopeListsJob` | **588** | 1004 B |
| `IndexUserJob` | **395** | 954 B |
| `RecComputeSimilarByTagsJob` | **169** | 1212 B |
| `IndexArtworkJob` | 40 | 918 B |
| `RecalculateRisingNovaCardsJob` | 10 | 987 B |
| `RebuildTrendingNovaCardsJob` | 3 | 977 B |
| `RankBuildListsJob` | 2 | 931 B |
| `RecBuildItemPairsFromFavouritesJob` | 1 | 1102 B |
- Oldest pushedAt: **2026-08-23 15:19** (~2.5 h)
- Newest: 17:48 (~75 s)
- Delayed: 1 `RankBuildScopeListsJob`
- Reserved: 0
- Attempts: 1206×0, 2×1
- Max payload: **1212 B** — not oversized job payloads
Net 60s: 1208 → 1213 (**accumulating**).
---
## 2. Throughput
Horizon metrics repo empty from this user (prefix split `skinbase_horizon` vs `skinbasenova_horizon`).
Evidence: backlog growing +5 net/min while 5 default-supervisor processes exist. Ranking jobs `timeout=300s`; rec tag jobs `timeout=600s`. A few long jobs starve `IndexUserJob` (30s, user-facing search).
---
## 3. Producers
| Producer | Job | Queue | Cadence | Volume |
| -------- | --- | ----- | ------- | ------ |
| `RankBuildListsJob` hourlyAt(15) | `RankBuildScopeListsJob` per category/content type | default | hourly fan-out | **~588/hour** still sitting |
| `UserStatsService::reindex` | `IndexUserJob` | **default** (bug; Scout config is `search`) | per stats change | 395 queued |
| Artwork indexers | `IndexArtworkJob` | **default** (same bug) | publish/trending/tags | 40 queued |
| Scheduler rec nightly | `RecComputeSimilarByTagsJob` batches | default (`RECOMMENDATIONS_QUEUE`) | nightly + leftover | 169 |
| Nova card crons | Rising/Trending rebuild | default | 15 min / hourly | 13 |
| `collections:dispatch-maintenance` hourlyAt(43) | Health/rec/dup | **collections** | hourly ~110 | **32k unconsumed** |
| Forum plugin | `AnalyzeForumPostJob` | **forum-moderation** | continuous | **552,888 unconsumed** |
---
## 4. Isolation
Move **IndexArtworkJob / IndexUserJob** to `search` (already on supervisor-default, currently idle). Same worker budget, better latency. Do **not** add processes.
Do **not** merge mail. Rec stays on default until ranking unique + M1 afterArtworkId is healthy.
**collections** and **forum-moderation** have **no Horizon supervisor** — isolation without a consumer is a dump. Do not add workers until RAM is freed (M6). Stop/limit dispatch first.
---
## 5. Predis 256 MB fatals (13 Aug, 18 Aug)
FPM error log (Skinbase pool, HTTP):
```text
Allowed memory size of 268435456 bytes exhausted
(tried to allocate 67108872 bytes)
vendor/predis/predis/src/Connection/StreamConnection.php:212
```
Two Redis values large enough to explode a 256 MB FPM worker when Predis reads the reply:
| Key | Type | Size | How it is read |
| --- | ---- | ---- | -------------- |
| `queues:forum-moderation` | list | **552,888 jobs ≈ 557 MB** | Horizon UI / any `LRANGE` of the queue |
| `skinbase:presence:online:index` | set | **2,480,292 members ≈ 272 MB** | `OnlineVisitorRepository::readIndexMembers()` → **`SMEMBERS`** then `all()` on moderation traffic page |
Index members have **TTL -1**; per-visitor keys TTL 300s. Expired visitors are never removed unless `all()` finishes — it cannot, so the set grows forever.
**Do not raise memory_limit.** Stop loading the full set/list in one Redis command.
Legacy prefix `skinbasenova-database-queues:*` still holds leftover lists (another ~11 MB default + forum copies).
---
## 6. Payload size
Default-queue jobs are **~1 KB**. Not the memory issue.
---
## 7. Largest Redis keys/prefixes (SCAN, names redacted)
| Group | Approx |
| ----- | -----: |
| `queues:forum-moderation` | 557 MB |
| `skinbase:presence:online:index` | 272 MB |
| `queues:collections` | 35 MB |
| `skinbasenova-database-queues:default` (legacy prefix) | 11 MB |
| `queues:forum-security` | 5.5 MB |
| Horizon recent/pending (legacy prefix) | ~3 MB |
Keyspace ~3.7k keys but a few lists/sets dominate `used_memory` (~879 MB).
---
## 8. search vs default supervisor
**Keep sharing supervisor-default.** Search workers are idle because index jobs were on `default`. After Index* → `search`, auto-balance can drain indexing without extra RAM.
---
## 9. Backlog classification
**Default: C+D** — hourly ranking fan-out + 10-minute rec jobs exceed 5 workers; index jobs wait behind them. Not a retry storm (attempts mostly 0).
**collections / forum-*: F** — orphan queues, growing for months (collections from 2026-04-30; forum still enqueuing today).
**failed_jobs:** 133 rows, **all today 14:37–14:56** (Horizon SIGTERM 14:36). Classes: RecCompute* `Typed property $afterArtworkId must not be accessed before initialization` on release `20260823-12*` — production rec job constructor mismatch. Enhance jobs: curl 127.0.0.1:8095 refused. **Do not retry.**
---
## 10. User-facing latency
`IndexUserJob` / `IndexArtworkJob` wait behind ranking/rec on default. Oldest index jobs share the 2.5 h window. Emails/notifications not on default (mail/notify queues empty). Uploads/vision queue empty.
---
## 11. Redis latency vs swap
ops/sec ~42, blocked clients 0, evicted 0. Slowlog parse inconclusive. Swapped Redis **will** hurt when someone reads a 557 MB list (that *is* the 256 MB fatal). Queue pop of 1 KB jobs is small.
---
## 12. Smallest fix (local; not deployed)
1. `IndexArtworkJob` + `IndexUserJob` → `scout.queue.queue` (`search`) + `ShouldBeUnique`.
2. `RankBuildScopeListsJob` `ShouldBeUnique` per scope so hourly fan-out cannot stack 588 copies.
3. Presence index: **SSCAN + cap 2000**, never `SMEMBERS` of 2.4M.
**Not in this deploy:** delete 552k forum jobs, add Horizon supervisors, raise PHP memory, raise Horizon processes.
**After RAM (M6):** optional `supervisor-collections` / forum workers, or stop those dispatchers.
---
## 13. Tests
`tests/Feature/Queue/DefaultQueueAssignmentTest.php`
---
## 14. Deploy
Ship the three job/presence changes. Reload Horizon so new job `onQueue` applies to **new** dispatches (existing 395 IndexUserJob stay on default until drained). No FPM memory_limit change. No Redis delete.
Rollback: revert those files; Horizon reload.
Risk: unique ranking skips a scope if a stale unique lock remains (`uniqueFor=360s`). Presence admin UI shows at most 2000 index members until a later prune job (write) is designed.
Production not modified.