Files
SkinbaseNova/docs/optimization-m7.1-queue-compatibility.md
klevze 8a80aae21e Ship production optimization M1-M12.5A: queues, metrics, HTTP observability, and vector search reliability.
Keep similar-ai from tripping the global circuit on a lone URL 502, clamp Qdrant search to 100, and add Server-Timing plus slow-request logging. Studio shared props, Academy S3 exists caching, heat chunking, and Redis/scheduler hygiene stay in this rollout.
2026-08-25 07:58:47 +02:00

9.3 KiB
Raw Permalink Blame History

M7.1 — Queue Compatibility & Backlog Containment

STATUS: COMPLETE + FINALIZED (local; production unchanged)

1. RecCompute serialization root cause

M1 added constructor-promoted private readonly ?int $afterArtworkId = null.

Constructor defaults are not applied on unserialize. PHP 8 leaves a typed property uninitialized if it is missing from the serialized payload. First access (handle, failed(), logging) throws:

Typed property App\Jobs\RecComputeSimilar*::$afterArtworkId must not be accessed before initialization

Old queued jobs (pre-M1, two constructor args) and Laravel SerializesModels (skips default-null properties on serialize) both omit the key.

Fix: class-level private ?int $afterArtworkId = null (not readonly promotion) + __unserialize calls isset restore. Getter afterArtworkId() never reads uninitialized state.

Cursor batching unchanged: WHERE id > afterArtworkId, ORDER BY id, LIMIT batch, no OFFSET.


2. Stale Tags jobs — worst case (live 2026-08-23 18:xx)

Live queues:default RecCompute* (not failed_jobs):

  • Behavior: 0 waiting
  • Hybrid: 0 waiting
  • Tags: 169 waiting, all artworkId=null (batch mode), all already have afterArtworkId set (cursors 1284–4517). These are overlapping remainder chains, not missing-property payloads.

What afterArtworkId=null would do: process first 200 public artworks ordered by id, then dispatch(null, 200, lastId) until the table is exhausted. One full rebuild ≈ ceil(N/200) jobs. With ~50k snapshot-eligible artworks that is ~250 jobs per chain.

If all 169 were null-cursor roots: 169 × ~250 ≈ 42,000 tag jobs (duplicate full rebuilds). That is the theoretical worst case.

Actual live worst case: they continue from mid-table. Example: 47 jobs with after=3709 each walk the remainder (~232 batches) → ~11,000 overlapping jobs from that cursor group alone, plus the other cursor groups. Still dangerous.

Do not clear queues:default wholesale. Operator command (not scheduled):

php artisan skinbase:purge-queued-rec-tags --queue=default --dry-run
php artisan skinbase:purge-queued-rec-tags --queue=default --execute

Removes waiting RecComputeSimilarByTagsJob only. Preserves ranking/index/other. Does not touch reserved or failed_jobs. Then tonight’s scheduler dispatches one clean Tags rebuild (afterArtworkId=null once).


3. Failed-job breakdown (complete, n=133, all queue=default)

M7 sampled 80 rows. Full table:

Class Count Cause
RecComputeSimilarHybridJob 66 $afterArtworkId uninitialized
RecComputeSimilarByBehaviorJob 42 same
RecComputeSimilarByTagsJob 10 same
Enhance\ProcessEnhanceJob 8 curl 127.0.0.1:8095 refused
App\Mail\AcademyAccessIssue 7 mail on default
Total 133

66+42+10+8+7=133. Do not retry.


4. Index* queue

IndexArtworkJob / IndexUserJob use config('scout.queue.queue') (default search).

ShouldBeUniqueUntilProcessing:

IndexUser IndexArtwork
uniqueId userId artworkId
uniqueFor 60s 120s

handle() loads the model from DB, so a dropped duplicate while queued still indexes current stats. After the worker starts, a later change can enqueue again (not lost). Dead-worker lock expires at uniqueFor.


5. Ranking dedupe

RankBuildScopeListsJob ShouldBeUnique, uniqueId = scopeType:scopeId.

uniqueFor=21600 (6 hours) from ranking.scope_job_unique_for / RANK_SCOPE_JOB_UNIQUE_FOR. 360s expired long before a 2.5h wait, so hourly fan-out stacked duplicates. 6h > current wait + 300s timeout; dead locks recover the same day. Different scopes stay independent.


6. Presence memory

SMEMBERS removed. SSCAN + cap traffic.online_visitors.index_read_limit (default 2000). Admin UI may under-count vs 2.48M stale members until prune.

Stale-member prune (design only — not executed)

  1. Detect: SSCAN members; GET skinbase:presence:online:{member}; if missing/expired JSON, member is stale.
  2. Batches: SSCAN COUNT 200; SREM up to 200 stale ids; sleep 50ms.
  3. Cost: 2.48M SSCAN + GET. At 200/batch ~12k round trips. Off-peak 5–15 min. Do not DEL the index key (would drop live members).
  4. Cadence: hourly withoutOverlapping, stop after N batches/run.

7. Orphan producers

Queue Consumer? Enabled? Useful if months late? Dispatch today? M7.1
collections No Horizon supervisor scheduler hourlyAt(43) Refresh is idempotent; delayed run OK yes until gate COLLECTIONS_V5_DISPATCH_ENABLED=false (default)
forum-moderation No Horizon supervisor gated No for months-old posts was yes dispatchAsyncScan() now returns without queueing when skinbase_ai_moderation.queue.dispatch_enabled is false. Topic create/edit and forum:ai-scan (async) are gated. forum:scan-posts and forum:ai-scan --sync stay in-process.
forum-security No gated stale in hours 5.5k forum:bot-scan / forum:firewall-scan async dispatch gated. --sync still runs. Scheduler omits those commands when dispatch is false.

Do not add Horizon workers (M6 RAM).


8. Existing orphan jobs (do not delete)

Queue Class Disposition
collections Health / recommendation / duplicate KEEP AND PROCESS LATER or SAFE TO DISCARD (recompute from current rows). Product: either.
forum-moderation AnalyzeForumPostJob × 552,888 SAFE TO DISCARD for age > 48h; NEEDS PRODUCT DECISION for last day. Analyzing May–August posts has no moderation value.
forum-security ~5.5k SAFE TO DISCARD (burst/firewall windows expired).

Later cleanup (not now): UNLINK queues:forum-moderation (async, non-blocking) after stopping producers; or LTRIM/RPOP in 5k batches. Never one giant DEL in a request.


9. Legacy prefix skinbasenova-database-queues:*

Current Laravel prefix is skinbase-database-. skinbasenova-database-* is leftover from an older REDIS_PREFIX / app name. Nothing in this codebase reads it. Horizon also has skinbasenova_horizon:* vs skinbase_horizon:*.

Reclaimable: ~11 MB default + ~9 MB forum + ~2.5 MB collections + Horizon recent lists (~3 MB) ≈ 25–30 MB (not the 557 MB live forum list).

Cleanup only after confirming no process uses skinbasenova-database- (HORIZON_PREFIX, REDIS_PREFIX on all PHP/Horizon). Then UNLINK those keys. Do not delete skinbase-database-queues:forum-moderation.


10. Files changed

app/Jobs/Concerns/RestoresAfterArtworkIdCursor.php
app/Jobs/RecComputeSimilarByTagsJob.php
app/Jobs/RecComputeSimilarByBehaviorJob.php
app/Jobs/RecComputeSimilarHybridJob.php
app/Jobs/IndexArtworkJob.php
app/Jobs/IndexUserJob.php
app/Jobs/RankBuildScopeListsJob.php
app/Services/Traffic/OnlineVisitorRepository.php
app/Services/CollectionBackgroundJobService.php
app/Console/Commands/DispatchCollectionMaintenanceCommand.php
routes/console.php
config/collections.php
config/forum_security.php
config/skinbase_ai_moderation.php
config/traffic.php
.env.example
phpunit.xml
tests/...
docs/optimization-m7.1-queue-compatibility.md

11. Tests

15 passed: RecCompute unserialize, cursor batching, Index* search queue, ranking uniqueId per scope, presence SSCAN/no SMEMBERS, collection dispatch gate.


12. Deploy sequence (later; not now)

  1. Ship this release (jobs, uniqueness, presence SSCAN, collection gate, purge command).
  2. Production .env:
    COLLECTIONS_V5_DISPATCH_ENABLED=false
    RANK_SCOPE_JOB_UNIQUE_FOR=21600
    
    Do not rely on FORUM_QUEUE_DISPATCH_ENABLED until the Forum plugin is changed.
  3. php artisan config:cache
  4. sudo supervisorctl restart skinbase-horizon
  5. Optional: sudo systemctl reload php8.4-fpm
  6. After Horizon is on the new code, purge stale Tags (operator, not cron):
    php artisan skinbase:purge-queued-rec-tags --queue=default --dry-run
    php artisan skinbase:purge-queued-rec-tags --queue=default --execute
    
  7. Confirm queues:default RecComputeSimilarByTagsJob count is 0. Leave ranking/index jobs.
  8. Tonight’s scheduler may dispatch one Tags rebuild.

Horizon restart: YES. Do not retry failed_jobs. Do not SMEMBERS presence index. Do not DEL forum-moderation.


13. Post-deploy verification

php artisan tinker --execute="echo json_encode([
  'default' => Redis::llen('queues:default'),
  'search' => Redis::llen('queues:search'),
  'collections' => Redis::llen('queues:collections'),
  'forum' => Redis::llen('queues:forum-moderation'),
]);"
# After some traffic, search LLEN should rise (Index* jobs), default ranking unique should stop stacking
# Do not retry failed_jobs
# Do not SMEMBERS skinbase:presence:online:index

14. Risks

  • 169 in-queue tag jobs become “start from id 0” batches; duplicate compute, extra load, no crash.
  • Ranking unique skips a scope for up to 360s if a lock is stuck.
  • Presence UI shows ≤2000 members until prune.
  • Forum enqueue may continue if the plugin ignores FORUM_QUEUE_DISPATCH_ENABLED.
  • Collection tests set dispatch enabled in phpunit; production default is false.