Files
SkinbaseNova/docs/optimization-m7.1-queue-compatibility.md
T
klevze 8a80aae21e Ship production optimization M1-M12.5A: queues, metrics, HTTP observability, and vector search reliability.
Keep similar-ai from tripping the global circuit on a lone URL 502, clamp Qdrant search to 100, and add Server-Timing plus slow-request logging. Studio shared props, Academy S3 exists caching, heat chunking, and Redis/scheduler hygiene stay in this rollout.
2026-08-25 07:58:47 +02:00

218 lines
9.3 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# M7.1 — Queue Compatibility & Backlog Containment
```text
STATUS: COMPLETE + FINALIZED (local; production unchanged)
```
---
## 1. RecCompute serialization root cause
M1 added constructor-promoted `private readonly ?int $afterArtworkId = null`.
**Constructor defaults are not applied on `unserialize`.** PHP 8 leaves a **typed property uninitialized** if it is missing from the serialized payload. First access (handle, failed(), logging) throws:
```text
Typed property App\Jobs\RecComputeSimilar*::$afterArtworkId must not be accessed before initialization
```
Old queued jobs (pre-M1, two constructor args) and Laravel `SerializesModels` (skips default-null properties on serialize) both omit the key.
**Fix:** class-level `private ?int $afterArtworkId = null` (not readonly promotion) + `__unserialize` calls `isset` restore. Getter `afterArtworkId()` never reads uninitialized state.
Cursor batching unchanged: `WHERE id > afterArtworkId`, `ORDER BY id`, `LIMIT batch`, no OFFSET.
---
## 2. Stale Tags jobs — worst case (live 2026-08-23 18:xx)
Live `queues:default` RecCompute* (not failed_jobs):
- **Behavior: 0 waiting**
- **Hybrid: 0 waiting**
- **Tags: 169 waiting**, all `artworkId=null` (batch mode), **all already have `afterArtworkId` set** (cursors 1284–4517). These are overlapping remainder chains, not missing-property payloads.
What `afterArtworkId=null` would do: process first 200 public artworks ordered by id, then `dispatch(null, 200, lastId)` until the table is exhausted. One full rebuild ≈ `ceil(N/200)` jobs. With ~50k snapshot-eligible artworks that is **~250 jobs per chain**.
If all 169 were null-cursor roots: **169 × ~250 ≈ 42,000 tag jobs** (duplicate full rebuilds). That is the theoretical worst case.
**Actual live worst case:** they continue from mid-table. Example: 47 jobs with `after=3709` each walk the remainder (~232 batches) → **~11,000 overlapping jobs from that cursor group alone**, plus the other cursor groups. Still dangerous.
**Do not clear `queues:default` wholesale.** Operator command (not scheduled):
```bash
php artisan skinbase:purge-queued-rec-tags --queue=default --dry-run
php artisan skinbase:purge-queued-rec-tags --queue=default --execute
```
Removes **waiting** `RecComputeSimilarByTagsJob` only. Preserves ranking/index/other. Does not touch reserved or `failed_jobs`. Then tonight’s scheduler dispatches **one** clean Tags rebuild (`afterArtworkId=null` once).
---
## 3. Failed-job breakdown (complete, n=133, all queue=default)
M7 sampled 80 rows. Full table:
| Class | Count | Cause |
| ----- | ----: | ----- |
| RecComputeSimilarHybridJob | **66** | `$afterArtworkId` uninitialized |
| RecComputeSimilarByBehaviorJob | **42** | same |
| RecComputeSimilarByTagsJob | **10** | same |
| Enhance\ProcessEnhanceJob | **8** | curl 127.0.0.1:8095 refused |
| App\Mail\AcademyAccessIssue | **7** | mail on default |
| **Total** | **133** | |
66+42+10+8+7=133. Do not retry.
---
## 4. Index* queue
`IndexArtworkJob` / `IndexUserJob` use `config('scout.queue.queue')` (default `search`).
`ShouldBeUniqueUntilProcessing`:
| | IndexUser | IndexArtwork |
| --- | --- | --- |
| uniqueId | userId | artworkId |
| uniqueFor | 60s | 120s |
handle() loads the model from DB, so a dropped duplicate while queued still indexes **current** stats. After the worker **starts**, a later change can enqueue again (not lost). Dead-worker lock expires at uniqueFor.
---
## 5. Ranking dedupe
`RankBuildScopeListsJob` `ShouldBeUnique`, `uniqueId = scopeType:scopeId`.
**uniqueFor=21600 (6 hours)** from `ranking.scope_job_unique_for` / `RANK_SCOPE_JOB_UNIQUE_FOR`. 360s expired long before a 2.5h wait, so hourly fan-out stacked duplicates. 6h > current wait + 300s timeout; dead locks recover the same day. Different scopes stay independent.
---
## 6. Presence memory
`SMEMBERS` removed. `SSCAN` + cap `traffic.online_visitors.index_read_limit` (default 2000). Admin UI may under-count vs 2.48M stale members until prune.
### Stale-member prune (design only — not executed)
1. Detect: SSCAN members; `GET skinbase:presence:online:{member}`; if missing/expired JSON, member is stale.
2. Batches: SSCAN COUNT 200; `SREM` up to 200 stale ids; sleep 50ms.
3. Cost: 2.48M SSCAN + GET. At 200/batch ~12k round trips. Off-peak 5–15 min. Do not `DEL` the index key (would drop live members).
4. Cadence: hourly `withoutOverlapping`, stop after N batches/run.
---
## 7. Orphan producers
| Queue | Consumer? | Enabled? | Useful if months late? | Dispatch today? | M7.1 |
| ----- | --------- | -------- | --------------------- | --------------- | ---- |
| collections | **No** Horizon supervisor | scheduler hourlyAt(43) | Refresh is idempotent; delayed run OK | yes until gate | **`COLLECTIONS_V5_DISPATCH_ENABLED=false`** (default) |
| forum-moderation | **No** Horizon supervisor | gated | **No** for months-old posts | was yes | **`dispatchAsyncScan()` now returns without queueing when `skinbase_ai_moderation.queue.dispatch_enabled` is false.** Topic create/edit and `forum:ai-scan` (async) are gated. `forum:scan-posts` and `forum:ai-scan --sync` stay in-process. |
| forum-security | **No** | gated | stale in hours | 5.5k | **`forum:bot-scan` / `forum:firewall-scan` async dispatch gated.** `--sync` still runs. Scheduler omits those commands when dispatch is false. |
Do not add Horizon workers (M6 RAM).
---
## 8. Existing orphan jobs (do not delete)
| Queue | Class | Disposition |
| ----- | ----- | ----------- |
| collections | Health / recommendation / duplicate | **KEEP AND PROCESS LATER** or **SAFE TO DISCARD** (recompute from current rows). Product: either. |
| forum-moderation | AnalyzeForumPostJob × 552,888 | **SAFE TO DISCARD** for age > 48h; **NEEDS PRODUCT DECISION** for last day. Analyzing May–August posts has no moderation value. |
| forum-security | ~5.5k | **SAFE TO DISCARD** (burst/firewall windows expired). |
Later cleanup (not now): `UNLINK queues:forum-moderation` (async, non-blocking) after stopping producers; or `LTRIM`/`RPOP` in 5k batches. Never one giant `DEL` in a request.
---
## 9. Legacy prefix `skinbasenova-database-queues:*`
Current Laravel prefix is `skinbase-database-`. `skinbasenova-database-*` is leftover from an older `REDIS_PREFIX` / app name. Nothing in this codebase reads it. Horizon also has `skinbasenova_horizon:*` vs `skinbase_horizon:*`.
Reclaimable: ~11 MB default + ~9 MB forum + ~2.5 MB collections + Horizon recent lists (~3 MB) ≈ **25–30 MB** (not the 557 MB live forum list).
Cleanup only after confirming no process uses `skinbasenova-database-` (`HORIZON_PREFIX`, `REDIS_PREFIX` on all PHP/Horizon). Then `UNLINK` those keys. Do not delete `skinbase-database-queues:forum-moderation`.
---
## 10. Files changed
```text
app/Jobs/Concerns/RestoresAfterArtworkIdCursor.php
app/Jobs/RecComputeSimilarByTagsJob.php
app/Jobs/RecComputeSimilarByBehaviorJob.php
app/Jobs/RecComputeSimilarHybridJob.php
app/Jobs/IndexArtworkJob.php
app/Jobs/IndexUserJob.php
app/Jobs/RankBuildScopeListsJob.php
app/Services/Traffic/OnlineVisitorRepository.php
app/Services/CollectionBackgroundJobService.php
app/Console/Commands/DispatchCollectionMaintenanceCommand.php
routes/console.php
config/collections.php
config/forum_security.php
config/skinbase_ai_moderation.php
config/traffic.php
.env.example
phpunit.xml
tests/...
docs/optimization-m7.1-queue-compatibility.md
```
---
## 11. Tests
15 passed: RecCompute unserialize, cursor batching, Index* search queue, ranking uniqueId per scope, presence SSCAN/no SMEMBERS, collection dispatch gate.
---
## 12. Deploy sequence (later; not now)
1. Ship this release (jobs, uniqueness, presence SSCAN, collection gate, purge command).
2. Production `.env`:
```text
COLLECTIONS_V5_DISPATCH_ENABLED=false
RANK_SCOPE_JOB_UNIQUE_FOR=21600
```
Do not rely on `FORUM_QUEUE_DISPATCH_ENABLED` until the Forum plugin is changed.
3. `php artisan config:cache`
4. `sudo supervisorctl restart skinbase-horizon`
5. Optional: `sudo systemctl reload php8.4-fpm`
6. **After Horizon is on the new code**, purge stale Tags (operator, not cron):
```bash
php artisan skinbase:purge-queued-rec-tags --queue=default --dry-run
php artisan skinbase:purge-queued-rec-tags --queue=default --execute
```
7. Confirm `queues:default` RecComputeSimilarByTagsJob count is 0. Leave ranking/index jobs.
8. Tonight’s scheduler may dispatch **one** Tags rebuild.
**Horizon restart: YES.** Do not retry `failed_jobs`. Do not `SMEMBERS` presence index. Do not `DEL` forum-moderation.
---
## 13. Post-deploy verification
```bash
php artisan tinker --execute="echo json_encode([
'default' => Redis::llen('queues:default'),
'search' => Redis::llen('queues:search'),
'collections' => Redis::llen('queues:collections'),
'forum' => Redis::llen('queues:forum-moderation'),
]);"
# After some traffic, search LLEN should rise (Index* jobs), default ranking unique should stop stacking
# Do not retry failed_jobs
# Do not SMEMBERS skinbase:presence:online:index
```
---
## 14. Risks
- 169 in-queue tag jobs become “start from id 0” batches; duplicate compute, extra load, no crash.
- Ranking unique skips a scope for up to 360s if a lock is stuck.
- Presence UI shows ≤2000 members until prune.
- Forum enqueue may continue if the plugin ignores `FORUM_QUEUE_DISPATCH_ENABLED`.
- Collection tests set dispatch enabled in phpunit; production default is false.