Keep similar-ai from tripping the global circuit on a lone URL 502, clamp Qdrant search to 100, and add Server-Timing plus slow-request logging. Studio shared props, Academy S3 exists caching, heat chunking, and Redis/scheduler hygiene stay in this rollout.
7.3 KiB
M7 — Default Queue Backlog & Redis/PHP Memory Failures
STATUS: COMPLETE (investigation + local smallest fixes; production unchanged)
WHEN: 2026-08-23 17:49 Europe/Ljubljana
No Horizon restart, Redis flush, job retry, memory_limit change, or worker-count change on server3.
Verdict
Default backlog class: C + D (producer > consumer; ranking/rec jobs block index jobs)
Default depth: 1208 → 1213 in 60s (net +5/min, growing)
Oldest default job: ~2.5 h
Search queue: 0 (idle workers exist on supervisor-default)
collections: 32,632 jobs, oldest 2026-04-30, NO Horizon consumer
forum-moderation: 552,888 AnalyzeForumPostJob, ~557 MB list, NO consumer
forum-security: 5,533 jobs, no consumer
presence index: SET 2,480,292 members, ~272 MB, TTL -1
Predis 256 MB fatals: loading those huge Redis replies over HTTP/FPM
Horizon workers: do NOT increase
search+default supervisor: KEEP together; put Index* jobs on search
1. queues:default contents (metadata only)
| Class | Count | Avg payload |
|---|---|---|
RankBuildScopeListsJob |
588 | 1004 B |
IndexUserJob |
395 | 954 B |
RecComputeSimilarByTagsJob |
169 | 1212 B |
IndexArtworkJob |
40 | 918 B |
RecalculateRisingNovaCardsJob |
10 | 987 B |
RebuildTrendingNovaCardsJob |
3 | 977 B |
RankBuildListsJob |
2 | 931 B |
RecBuildItemPairsFromFavouritesJob |
1 | 1102 B |
- Oldest pushedAt: 2026-08-23 15:19 (~2.5 h)
- Newest: 17:48 (~75 s)
- Delayed: 1
RankBuildScopeListsJob - Reserved: 0
- Attempts: 1206×0, 2×1
- Max payload: 1212 B — not oversized job payloads
Net 60s: 1208 → 1213 (accumulating).
2. Throughput
Horizon metrics repo empty from this user (prefix split skinbase_horizon vs skinbasenova_horizon).
Evidence: backlog growing +5 net/min while 5 default-supervisor processes exist. Ranking jobs timeout=300s; rec tag jobs timeout=600s. A few long jobs starve IndexUserJob (30s, user-facing search).
3. Producers
| Producer | Job | Queue | Cadence | Volume |
|---|---|---|---|---|
RankBuildListsJob hourlyAt(15) |
RankBuildScopeListsJob per category/content type |
default | hourly fan-out | ~588/hour still sitting |
UserStatsService::reindex |
IndexUserJob |
default (bug; Scout config is search) |
per stats change | 395 queued |
| Artwork indexers | IndexArtworkJob |
default (same bug) | publish/trending/tags | 40 queued |
| Scheduler rec nightly | RecComputeSimilarByTagsJob batches |
default (RECOMMENDATIONS_QUEUE) |
nightly + leftover | 169 |
| Nova card crons | Rising/Trending rebuild | default | 15 min / hourly | 13 |
collections:dispatch-maintenance hourlyAt(43) |
Health/rec/dup | collections | hourly ~110 | 32k unconsumed |
| Forum plugin | AnalyzeForumPostJob |
forum-moderation | continuous | 552,888 unconsumed |
4. Isolation
Move IndexArtworkJob / IndexUserJob to search (already on supervisor-default, currently idle). Same worker budget, better latency. Do not add processes.
Do not merge mail. Rec stays on default until ranking unique + M1 afterArtworkId is healthy.
collections and forum-moderation have no Horizon supervisor — isolation without a consumer is a dump. Do not add workers until RAM is freed (M6). Stop/limit dispatch first.
5. Predis 256 MB fatals (13 Aug, 18 Aug)
FPM error log (Skinbase pool, HTTP):
Allowed memory size of 268435456 bytes exhausted
(tried to allocate 67108872 bytes)
vendor/predis/predis/src/Connection/StreamConnection.php:212
Two Redis values large enough to explode a 256 MB FPM worker when Predis reads the reply:
| Key | Type | Size | How it is read |
|---|---|---|---|
queues:forum-moderation |
list | 552,888 jobs ≈ 557 MB | Horizon UI / any LRANGE of the queue |
skinbase:presence:online:index |
set | 2,480,292 members ≈ 272 MB | OnlineVisitorRepository::readIndexMembers() → SMEMBERS then all() on moderation traffic page |
Index members have TTL -1; per-visitor keys TTL 300s. Expired visitors are never removed unless all() finishes — it cannot, so the set grows forever.
Do not raise memory_limit. Stop loading the full set/list in one Redis command.
Legacy prefix skinbasenova-database-queues:* still holds leftover lists (another ~11 MB default + forum copies).
6. Payload size
Default-queue jobs are ~1 KB. Not the memory issue.
7. Largest Redis keys/prefixes (SCAN, names redacted)
| Group | Approx |
|---|---|
queues:forum-moderation |
557 MB |
skinbase:presence:online:index |
272 MB |
queues:collections |
35 MB |
skinbasenova-database-queues:default (legacy prefix) |
11 MB |
queues:forum-security |
5.5 MB |
| Horizon recent/pending (legacy prefix) | ~3 MB |
Keyspace ~3.7k keys but a few lists/sets dominate used_memory (~879 MB).
8. search vs default supervisor
Keep sharing supervisor-default. Search workers are idle because index jobs were on default. After Index* → search, auto-balance can drain indexing without extra RAM.
9. Backlog classification
Default: C+D — hourly ranking fan-out + 10-minute rec jobs exceed 5 workers; index jobs wait behind them. Not a retry storm (attempts mostly 0).
collections / forum-*: F — orphan queues, growing for months (collections from 2026-04-30; forum still enqueuing today).
failed_jobs: 133 rows, all today 14:37–14:56 (Horizon SIGTERM 14:36). Classes: RecCompute* Typed property $afterArtworkId must not be accessed before initialization on release 20260823-12* — production rec job constructor mismatch. Enhance jobs: curl 127.0.0.1:8095 refused. Do not retry.
10. User-facing latency
IndexUserJob / IndexArtworkJob wait behind ranking/rec on default. Oldest index jobs share the 2.5 h window. Emails/notifications not on default (mail/notify queues empty). Uploads/vision queue empty.
11. Redis latency vs swap
ops/sec ~42, blocked clients 0, evicted 0. Slowlog parse inconclusive. Swapped Redis will hurt when someone reads a 557 MB list (that is the 256 MB fatal). Queue pop of 1 KB jobs is small.
12. Smallest fix (local; not deployed)
IndexArtworkJob+IndexUserJob→scout.queue.queue(search) +ShouldBeUnique.RankBuildScopeListsJobShouldBeUniqueper scope so hourly fan-out cannot stack 588 copies.- Presence index: SSCAN + cap 2000, never
SMEMBERSof 2.4M.
Not in this deploy: delete 552k forum jobs, add Horizon supervisors, raise PHP memory, raise Horizon processes.
After RAM (M6): optional supervisor-collections / forum workers, or stop those dispatchers.
13. Tests
tests/Feature/Queue/DefaultQueueAssignmentTest.php
14. Deploy
Ship the three job/presence changes. Reload Horizon so new job onQueue applies to new dispatches (existing 395 IndexUserJob stay on default until drained). No FPM memory_limit change. No Redis delete.
Rollback: revert those files; Horizon reload.
Risk: unique ranking skips a scope if a stale unique lock remains (uniqueFor=360s). Presence admin UI shows at most 2000 index members until a later prune job (write) is designed.
Production not modified.