Files
SkinbaseNova/docs/optimization-m7-default-queue-backlog.md
klevze 8a80aae21e Ship production optimization M1-M12.5A: queues, metrics, HTTP observability, and vector search reliability.
Keep similar-ai from tripping the global circuit on a lone URL 502, clamp Qdrant search to 100, and add Server-Timing plus slow-request logging. Studio shared props, Academy S3 exists caching, heat chunking, and Redis/scheduler hygiene stay in this rollout.
2026-08-25 07:58:47 +02:00

187 lines
7.3 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# M7 — Default Queue Backlog & Redis/PHP Memory Failures
```text
STATUS: COMPLETE (investigation + local smallest fixes; production unchanged)
WHEN: 2026-08-23 17:49 Europe/Ljubljana
```
No Horizon restart, Redis flush, job retry, memory_limit change, or worker-count change on server3.
---
## Verdict
```text
Default backlog class: C + D (producer > consumer; ranking/rec jobs block index jobs)
Default depth: 1208 → 1213 in 60s (net +5/min, growing)
Oldest default job: ~2.5 h
Search queue: 0 (idle workers exist on supervisor-default)
collections: 32,632 jobs, oldest 2026-04-30, NO Horizon consumer
forum-moderation: 552,888 AnalyzeForumPostJob, ~557 MB list, NO consumer
forum-security: 5,533 jobs, no consumer
presence index: SET 2,480,292 members, ~272 MB, TTL -1
Predis 256 MB fatals: loading those huge Redis replies over HTTP/FPM
Horizon workers: do NOT increase
search+default supervisor: KEEP together; put Index* jobs on search
```
---
## 1. `queues:default` contents (metadata only)
| Class | Count | Avg payload |
| ----- | ----: | ----------: |
| `RankBuildScopeListsJob` | **588** | 1004 B |
| `IndexUserJob` | **395** | 954 B |
| `RecComputeSimilarByTagsJob` | **169** | 1212 B |
| `IndexArtworkJob` | 40 | 918 B |
| `RecalculateRisingNovaCardsJob` | 10 | 987 B |
| `RebuildTrendingNovaCardsJob` | 3 | 977 B |
| `RankBuildListsJob` | 2 | 931 B |
| `RecBuildItemPairsFromFavouritesJob` | 1 | 1102 B |
- Oldest pushedAt: **2026-08-23 15:19** (~2.5 h)
- Newest: 17:48 (~75 s)
- Delayed: 1 `RankBuildScopeListsJob`
- Reserved: 0
- Attempts: 1206×0, 2×1
- Max payload: **1212 B** — not oversized job payloads
Net 60s: 1208 → 1213 (**accumulating**).
---
## 2. Throughput
Horizon metrics repo empty from this user (prefix split `skinbase_horizon` vs `skinbasenova_horizon`).
Evidence: backlog growing +5 net/min while 5 default-supervisor processes exist. Ranking jobs `timeout=300s`; rec tag jobs `timeout=600s`. A few long jobs starve `IndexUserJob` (30s, user-facing search).
---
## 3. Producers
| Producer | Job | Queue | Cadence | Volume |
| -------- | --- | ----- | ------- | ------ |
| `RankBuildListsJob` hourlyAt(15) | `RankBuildScopeListsJob` per category/content type | default | hourly fan-out | **~588/hour** still sitting |
| `UserStatsService::reindex` | `IndexUserJob` | **default** (bug; Scout config is `search`) | per stats change | 395 queued |
| Artwork indexers | `IndexArtworkJob` | **default** (same bug) | publish/trending/tags | 40 queued |
| Scheduler rec nightly | `RecComputeSimilarByTagsJob` batches | default (`RECOMMENDATIONS_QUEUE`) | nightly + leftover | 169 |
| Nova card crons | Rising/Trending rebuild | default | 15 min / hourly | 13 |
| `collections:dispatch-maintenance` hourlyAt(43) | Health/rec/dup | **collections** | hourly ~110 | **32k unconsumed** |
| Forum plugin | `AnalyzeForumPostJob` | **forum-moderation** | continuous | **552,888 unconsumed** |
---
## 4. Isolation
Move **IndexArtworkJob / IndexUserJob** to `search` (already on supervisor-default, currently idle). Same worker budget, better latency. Do **not** add processes.
Do **not** merge mail. Rec stays on default until ranking unique + M1 afterArtworkId is healthy.
**collections** and **forum-moderation** have **no Horizon supervisor** — isolation without a consumer is a dump. Do not add workers until RAM is freed (M6). Stop/limit dispatch first.
---
## 5. Predis 256 MB fatals (13 Aug, 18 Aug)
FPM error log (Skinbase pool, HTTP):
```text
Allowed memory size of 268435456 bytes exhausted
(tried to allocate 67108872 bytes)
vendor/predis/predis/src/Connection/StreamConnection.php:212
```
Two Redis values large enough to explode a 256 MB FPM worker when Predis reads the reply:
| Key | Type | Size | How it is read |
| --- | ---- | ---- | -------------- |
| `queues:forum-moderation` | list | **552,888 jobs ≈ 557 MB** | Horizon UI / any `LRANGE` of the queue |
| `skinbase:presence:online:index` | set | **2,480,292 members ≈ 272 MB** | `OnlineVisitorRepository::readIndexMembers()` → **`SMEMBERS`** then `all()` on moderation traffic page |
Index members have **TTL -1**; per-visitor keys TTL 300s. Expired visitors are never removed unless `all()` finishes — it cannot, so the set grows forever.
**Do not raise memory_limit.** Stop loading the full set/list in one Redis command.
Legacy prefix `skinbasenova-database-queues:*` still holds leftover lists (another ~11 MB default + forum copies).
---
## 6. Payload size
Default-queue jobs are **~1 KB**. Not the memory issue.
---
## 7. Largest Redis keys/prefixes (SCAN, names redacted)
| Group | Approx |
| ----- | -----: |
| `queues:forum-moderation` | 557 MB |
| `skinbase:presence:online:index` | 272 MB |
| `queues:collections` | 35 MB |
| `skinbasenova-database-queues:default` (legacy prefix) | 11 MB |
| `queues:forum-security` | 5.5 MB |
| Horizon recent/pending (legacy prefix) | ~3 MB |
Keyspace ~3.7k keys but a few lists/sets dominate `used_memory` (~879 MB).
---
## 8. search vs default supervisor
**Keep sharing supervisor-default.** Search workers are idle because index jobs were on `default`. After Index* → `search`, auto-balance can drain indexing without extra RAM.
---
## 9. Backlog classification
**Default: C+D** — hourly ranking fan-out + 10-minute rec jobs exceed 5 workers; index jobs wait behind them. Not a retry storm (attempts mostly 0).
**collections / forum-*: F** — orphan queues, growing for months (collections from 2026-04-30; forum still enqueuing today).
**failed_jobs:** 133 rows, **all today 14:37–14:56** (Horizon SIGTERM 14:36). Classes: RecCompute* `Typed property $afterArtworkId must not be accessed before initialization` on release `20260823-12*` — production rec job constructor mismatch. Enhance jobs: curl 127.0.0.1:8095 refused. **Do not retry.**
---
## 10. User-facing latency
`IndexUserJob` / `IndexArtworkJob` wait behind ranking/rec on default. Oldest index jobs share the 2.5 h window. Emails/notifications not on default (mail/notify queues empty). Uploads/vision queue empty.
---
## 11. Redis latency vs swap
ops/sec ~42, blocked clients 0, evicted 0. Slowlog parse inconclusive. Swapped Redis **will** hurt when someone reads a 557 MB list (that *is* the 256 MB fatal). Queue pop of 1 KB jobs is small.
---
## 12. Smallest fix (local; not deployed)
1. `IndexArtworkJob` + `IndexUserJob` → `scout.queue.queue` (`search`) + `ShouldBeUnique`.
2. `RankBuildScopeListsJob` `ShouldBeUnique` per scope so hourly fan-out cannot stack 588 copies.
3. Presence index: **SSCAN + cap 2000**, never `SMEMBERS` of 2.4M.
**Not in this deploy:** delete 552k forum jobs, add Horizon supervisors, raise PHP memory, raise Horizon processes.
**After RAM (M6):** optional `supervisor-collections` / forum workers, or stop those dispatchers.
---
## 13. Tests
`tests/Feature/Queue/DefaultQueueAssignmentTest.php`
---
## 14. Deploy
Ship the three job/presence changes. Reload Horizon so new job `onQueue` applies to **new** dispatches (existing 395 IndexUserJob stay on default until drained). No FPM memory_limit change. No Redis delete.
Rollback: revert those files; Horizon reload.
Risk: unique ranking skips a scope if a stale unique lock remains (`uniqueFor=360s`). Presence admin UI shows at most 2000 index members until a later prune job (write) is designed.
Production not modified.