Files
SkinbaseNova/docs/optimization-m7-default-queue-backlog.md
T
klevze 8a80aae21e Ship production optimization M1-M12.5A: queues, metrics, HTTP observability, and vector search reliability.
Keep similar-ai from tripping the global circuit on a lone URL 502, clamp Qdrant search to 100, and add Server-Timing plus slow-request logging. Studio shared props, Academy S3 exists caching, heat chunking, and Redis/scheduler hygiene stay in this rollout.
2026-08-25 07:58:47 +02:00

7.3 KiB
Raw Blame History

M7 — Default Queue Backlog & Redis/PHP Memory Failures

STATUS: COMPLETE (investigation + local smallest fixes; production unchanged)
WHEN: 2026-08-23 17:49 Europe/Ljubljana

No Horizon restart, Redis flush, job retry, memory_limit change, or worker-count change on server3.


Verdict

Default backlog class:     C + D (producer > consumer; ranking/rec jobs block index jobs)
Default depth:             1208 → 1213 in 60s (net +5/min, growing)
Oldest default job:        ~2.5 h
Search queue:              0 (idle workers exist on supervisor-default)
collections:               32,632 jobs, oldest 2026-04-30, NO Horizon consumer
forum-moderation:          552,888 AnalyzeForumPostJob, ~557 MB list, NO consumer
forum-security:            5,533 jobs, no consumer
presence index:            SET 2,480,292 members, ~272 MB, TTL -1
Predis 256 MB fatals:      loading those huge Redis replies over HTTP/FPM
Horizon workers:           do NOT increase
search+default supervisor: KEEP together; put Index* jobs on search

1. queues:default contents (metadata only)

Class Count Avg payload
RankBuildScopeListsJob 588 1004 B
IndexUserJob 395 954 B
RecComputeSimilarByTagsJob 169 1212 B
IndexArtworkJob 40 918 B
RecalculateRisingNovaCardsJob 10 987 B
RebuildTrendingNovaCardsJob 3 977 B
RankBuildListsJob 2 931 B
RecBuildItemPairsFromFavouritesJob 1 1102 B
  • Oldest pushedAt: 2026-08-23 15:19 (~2.5 h)
  • Newest: 17:48 (~75 s)
  • Delayed: 1 RankBuildScopeListsJob
  • Reserved: 0
  • Attempts: 1206×0, 2×1
  • Max payload: 1212 B — not oversized job payloads

Net 60s: 1208 → 1213 (accumulating).


2. Throughput

Horizon metrics repo empty from this user (prefix split skinbase_horizon vs skinbasenova_horizon).

Evidence: backlog growing +5 net/min while 5 default-supervisor processes exist. Ranking jobs timeout=300s; rec tag jobs timeout=600s. A few long jobs starve IndexUserJob (30s, user-facing search).


3. Producers

Producer Job Queue Cadence Volume
RankBuildListsJob hourlyAt(15) RankBuildScopeListsJob per category/content type default hourly fan-out ~588/hour still sitting
UserStatsService::reindex IndexUserJob default (bug; Scout config is search) per stats change 395 queued
Artwork indexers IndexArtworkJob default (same bug) publish/trending/tags 40 queued
Scheduler rec nightly RecComputeSimilarByTagsJob batches default (RECOMMENDATIONS_QUEUE) nightly + leftover 169
Nova card crons Rising/Trending rebuild default 15 min / hourly 13
collections:dispatch-maintenance hourlyAt(43) Health/rec/dup collections hourly ~110 32k unconsumed
Forum plugin AnalyzeForumPostJob forum-moderation continuous 552,888 unconsumed

4. Isolation

Move IndexArtworkJob / IndexUserJob to search (already on supervisor-default, currently idle). Same worker budget, better latency. Do not add processes.

Do not merge mail. Rec stays on default until ranking unique + M1 afterArtworkId is healthy.

collections and forum-moderation have no Horizon supervisor — isolation without a consumer is a dump. Do not add workers until RAM is freed (M6). Stop/limit dispatch first.


5. Predis 256 MB fatals (13 Aug, 18 Aug)

FPM error log (Skinbase pool, HTTP):

Allowed memory size of 268435456 bytes exhausted
(tried to allocate 67108872 bytes)
vendor/predis/predis/src/Connection/StreamConnection.php:212

Two Redis values large enough to explode a 256 MB FPM worker when Predis reads the reply:

Key Type Size How it is read
queues:forum-moderation list 552,888 jobs ≈ 557 MB Horizon UI / any LRANGE of the queue
skinbase:presence:online:index set 2,480,292 members ≈ 272 MB OnlineVisitorRepository::readIndexMembers() → SMEMBERS then all() on moderation traffic page

Index members have TTL -1; per-visitor keys TTL 300s. Expired visitors are never removed unless all() finishes — it cannot, so the set grows forever.

Do not raise memory_limit. Stop loading the full set/list in one Redis command.

Legacy prefix skinbasenova-database-queues:* still holds leftover lists (another ~11 MB default + forum copies).


6. Payload size

Default-queue jobs are ~1 KB. Not the memory issue.


7. Largest Redis keys/prefixes (SCAN, names redacted)

Group Approx
queues:forum-moderation 557 MB
skinbase:presence:online:index 272 MB
queues:collections 35 MB
skinbasenova-database-queues:default (legacy prefix) 11 MB
queues:forum-security 5.5 MB
Horizon recent/pending (legacy prefix) ~3 MB

Keyspace ~3.7k keys but a few lists/sets dominate used_memory (~879 MB).


8. search vs default supervisor

Keep sharing supervisor-default. Search workers are idle because index jobs were on default. After Index* → search, auto-balance can drain indexing without extra RAM.


9. Backlog classification

Default: C+D — hourly ranking fan-out + 10-minute rec jobs exceed 5 workers; index jobs wait behind them. Not a retry storm (attempts mostly 0).

collections / forum-*: F — orphan queues, growing for months (collections from 2026-04-30; forum still enqueuing today).

failed_jobs: 133 rows, all today 14:37–14:56 (Horizon SIGTERM 14:36). Classes: RecCompute* Typed property $afterArtworkId must not be accessed before initialization on release 20260823-12* — production rec job constructor mismatch. Enhance jobs: curl 127.0.0.1:8095 refused. Do not retry.


10. User-facing latency

IndexUserJob / IndexArtworkJob wait behind ranking/rec on default. Oldest index jobs share the 2.5 h window. Emails/notifications not on default (mail/notify queues empty). Uploads/vision queue empty.


11. Redis latency vs swap

ops/sec ~42, blocked clients 0, evicted 0. Slowlog parse inconclusive. Swapped Redis will hurt when someone reads a 557 MB list (that is the 256 MB fatal). Queue pop of 1 KB jobs is small.


12. Smallest fix (local; not deployed)

  1. IndexArtworkJob + IndexUserJob → scout.queue.queue (search) + ShouldBeUnique.
  2. RankBuildScopeListsJob ShouldBeUnique per scope so hourly fan-out cannot stack 588 copies.
  3. Presence index: SSCAN + cap 2000, never SMEMBERS of 2.4M.

Not in this deploy: delete 552k forum jobs, add Horizon supervisors, raise PHP memory, raise Horizon processes.

After RAM (M6): optional supervisor-collections / forum workers, or stop those dispatchers.


13. Tests

tests/Feature/Queue/DefaultQueueAssignmentTest.php


14. Deploy

Ship the three job/presence changes. Reload Horizon so new job onQueue applies to new dispatches (existing 395 IndexUserJob stay on default until drained). No FPM memory_limit change. No Redis delete.

Rollback: revert those files; Horizon reload.

Risk: unique ranking skips a scope if a stale unique lock remains (uniqueFor=360s). Presence admin UI shows at most 2000 index members until a later prune job (write) is designed.

Production not modified.