Keep similar-ai from tripping the global circuit on a lone URL 502, clamp Qdrant search to 100, and add Server-Timing plus slow-request logging. Studio shared props, Academy S3 exists caching, heat chunking, and Redis/scheduler hygiene stay in this rollout.
9.3 KiB
M7.1 — Queue Compatibility & Backlog Containment
STATUS: COMPLETE + FINALIZED (local; production unchanged)
1. RecCompute serialization root cause
M1 added constructor-promoted private readonly ?int $afterArtworkId = null.
Constructor defaults are not applied on unserialize. PHP 8 leaves a typed property uninitialized if it is missing from the serialized payload. First access (handle, failed(), logging) throws:
Typed property App\Jobs\RecComputeSimilar*::$afterArtworkId must not be accessed before initialization
Old queued jobs (pre-M1, two constructor args) and Laravel SerializesModels (skips default-null properties on serialize) both omit the key.
Fix: class-level private ?int $afterArtworkId = null (not readonly promotion) + __unserialize calls isset restore. Getter afterArtworkId() never reads uninitialized state.
Cursor batching unchanged: WHERE id > afterArtworkId, ORDER BY id, LIMIT batch, no OFFSET.
2. Stale Tags jobs — worst case (live 2026-08-23 18:xx)
Live queues:default RecCompute* (not failed_jobs):
- Behavior: 0 waiting
- Hybrid: 0 waiting
- Tags: 169 waiting, all
artworkId=null(batch mode), all already haveafterArtworkIdset (cursors 1284–4517). These are overlapping remainder chains, not missing-property payloads.
What afterArtworkId=null would do: process first 200 public artworks ordered by id, then dispatch(null, 200, lastId) until the table is exhausted. One full rebuild ≈ ceil(N/200) jobs. With ~50k snapshot-eligible artworks that is ~250 jobs per chain.
If all 169 were null-cursor roots: 169 × ~250 ≈ 42,000 tag jobs (duplicate full rebuilds). That is the theoretical worst case.
Actual live worst case: they continue from mid-table. Example: 47 jobs with after=3709 each walk the remainder (~232 batches) → ~11,000 overlapping jobs from that cursor group alone, plus the other cursor groups. Still dangerous.
Do not clear queues:default wholesale. Operator command (not scheduled):
php artisan skinbase:purge-queued-rec-tags --queue=default --dry-run
php artisan skinbase:purge-queued-rec-tags --queue=default --execute
Removes waiting RecComputeSimilarByTagsJob only. Preserves ranking/index/other. Does not touch reserved or failed_jobs. Then tonight’s scheduler dispatches one clean Tags rebuild (afterArtworkId=null once).
3. Failed-job breakdown (complete, n=133, all queue=default)
M7 sampled 80 rows. Full table:
| Class | Count | Cause |
|---|---|---|
| RecComputeSimilarHybridJob | 66 | $afterArtworkId uninitialized |
| RecComputeSimilarByBehaviorJob | 42 | same |
| RecComputeSimilarByTagsJob | 10 | same |
| Enhance\ProcessEnhanceJob | 8 | curl 127.0.0.1:8095 refused |
| App\Mail\AcademyAccessIssue | 7 | mail on default |
| Total | 133 |
66+42+10+8+7=133. Do not retry.
4. Index* queue
IndexArtworkJob / IndexUserJob use config('scout.queue.queue') (default search).
ShouldBeUniqueUntilProcessing:
| IndexUser | IndexArtwork | |
|---|---|---|
| uniqueId | userId | artworkId |
| uniqueFor | 60s | 120s |
handle() loads the model from DB, so a dropped duplicate while queued still indexes current stats. After the worker starts, a later change can enqueue again (not lost). Dead-worker lock expires at uniqueFor.
5. Ranking dedupe
RankBuildScopeListsJob ShouldBeUnique, uniqueId = scopeType:scopeId.
uniqueFor=21600 (6 hours) from ranking.scope_job_unique_for / RANK_SCOPE_JOB_UNIQUE_FOR. 360s expired long before a 2.5h wait, so hourly fan-out stacked duplicates. 6h > current wait + 300s timeout; dead locks recover the same day. Different scopes stay independent.
6. Presence memory
SMEMBERS removed. SSCAN + cap traffic.online_visitors.index_read_limit (default 2000). Admin UI may under-count vs 2.48M stale members until prune.
Stale-member prune (design only — not executed)
- Detect: SSCAN members;
GET skinbase:presence:online:{member}; if missing/expired JSON, member is stale. - Batches: SSCAN COUNT 200;
SREMup to 200 stale ids; sleep 50ms. - Cost: 2.48M SSCAN + GET. At 200/batch ~12k round trips. Off-peak 5–15 min. Do not
DELthe index key (would drop live members). - Cadence: hourly
withoutOverlapping, stop after N batches/run.
7. Orphan producers
| Queue | Consumer? | Enabled? | Useful if months late? | Dispatch today? | M7.1 |
|---|---|---|---|---|---|
| collections | No Horizon supervisor | scheduler hourlyAt(43) | Refresh is idempotent; delayed run OK | yes until gate | COLLECTIONS_V5_DISPATCH_ENABLED=false (default) |
| forum-moderation | No Horizon supervisor | gated | No for months-old posts | was yes | dispatchAsyncScan() now returns without queueing when skinbase_ai_moderation.queue.dispatch_enabled is false. Topic create/edit and forum:ai-scan (async) are gated. forum:scan-posts and forum:ai-scan --sync stay in-process. |
| forum-security | No | gated | stale in hours | 5.5k | forum:bot-scan / forum:firewall-scan async dispatch gated. --sync still runs. Scheduler omits those commands when dispatch is false. |
Do not add Horizon workers (M6 RAM).
8. Existing orphan jobs (do not delete)
| Queue | Class | Disposition |
|---|---|---|
| collections | Health / recommendation / duplicate | KEEP AND PROCESS LATER or SAFE TO DISCARD (recompute from current rows). Product: either. |
| forum-moderation | AnalyzeForumPostJob × 552,888 | SAFE TO DISCARD for age > 48h; NEEDS PRODUCT DECISION for last day. Analyzing May–August posts has no moderation value. |
| forum-security | ~5.5k | SAFE TO DISCARD (burst/firewall windows expired). |
Later cleanup (not now): UNLINK queues:forum-moderation (async, non-blocking) after stopping producers; or LTRIM/RPOP in 5k batches. Never one giant DEL in a request.
9. Legacy prefix skinbasenova-database-queues:*
Current Laravel prefix is skinbase-database-. skinbasenova-database-* is leftover from an older REDIS_PREFIX / app name. Nothing in this codebase reads it. Horizon also has skinbasenova_horizon:* vs skinbase_horizon:*.
Reclaimable: ~11 MB default + ~9 MB forum + ~2.5 MB collections + Horizon recent lists (~3 MB) ≈ 25–30 MB (not the 557 MB live forum list).
Cleanup only after confirming no process uses skinbasenova-database- (HORIZON_PREFIX, REDIS_PREFIX on all PHP/Horizon). Then UNLINK those keys. Do not delete skinbase-database-queues:forum-moderation.
10. Files changed
app/Jobs/Concerns/RestoresAfterArtworkIdCursor.php
app/Jobs/RecComputeSimilarByTagsJob.php
app/Jobs/RecComputeSimilarByBehaviorJob.php
app/Jobs/RecComputeSimilarHybridJob.php
app/Jobs/IndexArtworkJob.php
app/Jobs/IndexUserJob.php
app/Jobs/RankBuildScopeListsJob.php
app/Services/Traffic/OnlineVisitorRepository.php
app/Services/CollectionBackgroundJobService.php
app/Console/Commands/DispatchCollectionMaintenanceCommand.php
routes/console.php
config/collections.php
config/forum_security.php
config/skinbase_ai_moderation.php
config/traffic.php
.env.example
phpunit.xml
tests/...
docs/optimization-m7.1-queue-compatibility.md
11. Tests
15 passed: RecCompute unserialize, cursor batching, Index* search queue, ranking uniqueId per scope, presence SSCAN/no SMEMBERS, collection dispatch gate.
12. Deploy sequence (later; not now)
- Ship this release (jobs, uniqueness, presence SSCAN, collection gate, purge command).
- Production
.env:Do not rely onCOLLECTIONS_V5_DISPATCH_ENABLED=false RANK_SCOPE_JOB_UNIQUE_FOR=21600FORUM_QUEUE_DISPATCH_ENABLEDuntil the Forum plugin is changed. php artisan config:cachesudo supervisorctl restart skinbase-horizon- Optional:
sudo systemctl reload php8.4-fpm - After Horizon is on the new code, purge stale Tags (operator, not cron):
php artisan skinbase:purge-queued-rec-tags --queue=default --dry-run php artisan skinbase:purge-queued-rec-tags --queue=default --execute - Confirm
queues:defaultRecComputeSimilarByTagsJob count is 0. Leave ranking/index jobs. - Tonight’s scheduler may dispatch one Tags rebuild.
Horizon restart: YES. Do not retry failed_jobs. Do not SMEMBERS presence index. Do not DEL forum-moderation.
13. Post-deploy verification
php artisan tinker --execute="echo json_encode([
'default' => Redis::llen('queues:default'),
'search' => Redis::llen('queues:search'),
'collections' => Redis::llen('queues:collections'),
'forum' => Redis::llen('queues:forum-moderation'),
]);"
# After some traffic, search LLEN should rise (Index* jobs), default ranking unique should stop stacking
# Do not retry failed_jobs
# Do not SMEMBERS skinbase:presence:online:index
14. Risks
- 169 in-queue tag jobs become “start from id 0” batches; duplicate compute, extra load, no crash.
- Ranking unique skips a scope for up to 360s if a lock is stuck.
- Presence UI shows ≤2000 members until prune.
- Forum enqueue may continue if the plugin ignores
FORUM_QUEUE_DISPATCH_ENABLED. - Collection tests set dispatch enabled in phpunit; production default is false.