Ship production optimization M1-M12.5A: queues, metrics, HTTP observability, and vector search reliability.

Keep similar-ai from tripping the global circuit on a lone URL 502, clamp Qdrant search to 100, and add Server-Timing plus slow-request logging. Studio shared props, Academy S3 exists caching, heat chunking, and Redis/scheduler hygiene stay in this rollout.
This commit is contained in:
2026-08-25 07:58:47 +02:00
parent f52879edbb
commit 8a80aae21e
114 changed files with 9293 additions and 284 deletions
@@ -0,0 +1,217 @@
# M7.1 — Queue Compatibility & Backlog Containment
```text
STATUS: COMPLETE + FINALIZED (local; production unchanged)
```
---
## 1. RecCompute serialization root cause
M1 added constructor-promoted `private readonly ?int $afterArtworkId = null`.
**Constructor defaults are not applied on `unserialize`.** PHP 8 leaves a **typed property uninitialized** if it is missing from the serialized payload. First access (handle, failed(), logging) throws:
```text
Typed property App\Jobs\RecComputeSimilar*::$afterArtworkId must not be accessed before initialization
```
Old queued jobs (pre-M1, two constructor args) and Laravel `SerializesModels` (skips default-null properties on serialize) both omit the key.
**Fix:** class-level `private ?int $afterArtworkId = null` (not readonly promotion) + `__unserialize` calls `isset` restore. Getter `afterArtworkId()` never reads uninitialized state.
Cursor batching unchanged: `WHERE id > afterArtworkId`, `ORDER BY id`, `LIMIT batch`, no OFFSET.
---
## 2. Stale Tags jobs — worst case (live 2026-08-23 18:xx)
Live `queues:default` RecCompute* (not failed_jobs):
- **Behavior: 0 waiting**
- **Hybrid: 0 waiting**
- **Tags: 169 waiting**, all `artworkId=null` (batch mode), **all already have `afterArtworkId` set** (cursors 1284–4517). These are overlapping remainder chains, not missing-property payloads.
What `afterArtworkId=null` would do: process first 200 public artworks ordered by id, then `dispatch(null, 200, lastId)` until the table is exhausted. One full rebuild ≈ `ceil(N/200)` jobs. With ~50k snapshot-eligible artworks that is **~250 jobs per chain**.
If all 169 were null-cursor roots: **169 × ~250 ≈ 42,000 tag jobs** (duplicate full rebuilds). That is the theoretical worst case.
**Actual live worst case:** they continue from mid-table. Example: 47 jobs with `after=3709` each walk the remainder (~232 batches) → **~11,000 overlapping jobs from that cursor group alone**, plus the other cursor groups. Still dangerous.
**Do not clear `queues:default` wholesale.** Operator command (not scheduled):
```bash
php artisan skinbase:purge-queued-rec-tags --queue=default --dry-run
php artisan skinbase:purge-queued-rec-tags --queue=default --execute
```
Removes **waiting** `RecComputeSimilarByTagsJob` only. Preserves ranking/index/other. Does not touch reserved or `failed_jobs`. Then tonight’s scheduler dispatches **one** clean Tags rebuild (`afterArtworkId=null` once).
---
## 3. Failed-job breakdown (complete, n=133, all queue=default)
M7 sampled 80 rows. Full table:
| Class | Count | Cause |
| ----- | ----: | ----- |
| RecComputeSimilarHybridJob | **66** | `$afterArtworkId` uninitialized |
| RecComputeSimilarByBehaviorJob | **42** | same |
| RecComputeSimilarByTagsJob | **10** | same |
| Enhance\ProcessEnhanceJob | **8** | curl 127.0.0.1:8095 refused |
| App\Mail\AcademyAccessIssue | **7** | mail on default |
| **Total** | **133** | |
66+42+10+8+7=133. Do not retry.
---
## 4. Index* queue
`IndexArtworkJob` / `IndexUserJob` use `config('scout.queue.queue')` (default `search`).
`ShouldBeUniqueUntilProcessing`:
| | IndexUser | IndexArtwork |
| --- | --- | --- |
| uniqueId | userId | artworkId |
| uniqueFor | 60s | 120s |
handle() loads the model from DB, so a dropped duplicate while queued still indexes **current** stats. After the worker **starts**, a later change can enqueue again (not lost). Dead-worker lock expires at uniqueFor.
---
## 5. Ranking dedupe
`RankBuildScopeListsJob` `ShouldBeUnique`, `uniqueId = scopeType:scopeId`.
**uniqueFor=21600 (6 hours)** from `ranking.scope_job_unique_for` / `RANK_SCOPE_JOB_UNIQUE_FOR`. 360s expired long before a 2.5h wait, so hourly fan-out stacked duplicates. 6h > current wait + 300s timeout; dead locks recover the same day. Different scopes stay independent.
---
## 6. Presence memory
`SMEMBERS` removed. `SSCAN` + cap `traffic.online_visitors.index_read_limit` (default 2000). Admin UI may under-count vs 2.48M stale members until prune.
### Stale-member prune (design only — not executed)
1. Detect: SSCAN members; `GET skinbase:presence:online:{member}`; if missing/expired JSON, member is stale.
2. Batches: SSCAN COUNT 200; `SREM` up to 200 stale ids; sleep 50ms.
3. Cost: 2.48M SSCAN + GET. At 200/batch ~12k round trips. Off-peak 5–15 min. Do not `DEL` the index key (would drop live members).
4. Cadence: hourly `withoutOverlapping`, stop after N batches/run.
---
## 7. Orphan producers
| Queue | Consumer? | Enabled? | Useful if months late? | Dispatch today? | M7.1 |
| ----- | --------- | -------- | --------------------- | --------------- | ---- |
| collections | **No** Horizon supervisor | scheduler hourlyAt(43) | Refresh is idempotent; delayed run OK | yes until gate | **`COLLECTIONS_V5_DISPATCH_ENABLED=false`** (default) |
| forum-moderation | **No** Horizon supervisor | gated | **No** for months-old posts | was yes | **`dispatchAsyncScan()` now returns without queueing when `skinbase_ai_moderation.queue.dispatch_enabled` is false.** Topic create/edit and `forum:ai-scan` (async) are gated. `forum:scan-posts` and `forum:ai-scan --sync` stay in-process. |
| forum-security | **No** | gated | stale in hours | 5.5k | **`forum:bot-scan` / `forum:firewall-scan` async dispatch gated.** `--sync` still runs. Scheduler omits those commands when dispatch is false. |
Do not add Horizon workers (M6 RAM).
---
## 8. Existing orphan jobs (do not delete)
| Queue | Class | Disposition |
| ----- | ----- | ----------- |
| collections | Health / recommendation / duplicate | **KEEP AND PROCESS LATER** or **SAFE TO DISCARD** (recompute from current rows). Product: either. |
| forum-moderation | AnalyzeForumPostJob × 552,888 | **SAFE TO DISCARD** for age > 48h; **NEEDS PRODUCT DECISION** for last day. Analyzing May–August posts has no moderation value. |
| forum-security | ~5.5k | **SAFE TO DISCARD** (burst/firewall windows expired). |
Later cleanup (not now): `UNLINK queues:forum-moderation` (async, non-blocking) after stopping producers; or `LTRIM`/`RPOP` in 5k batches. Never one giant `DEL` in a request.
---
## 9. Legacy prefix `skinbasenova-database-queues:*`
Current Laravel prefix is `skinbase-database-`. `skinbasenova-database-*` is leftover from an older `REDIS_PREFIX` / app name. Nothing in this codebase reads it. Horizon also has `skinbasenova_horizon:*` vs `skinbase_horizon:*`.
Reclaimable: ~11 MB default + ~9 MB forum + ~2.5 MB collections + Horizon recent lists (~3 MB) ≈ **25–30 MB** (not the 557 MB live forum list).
Cleanup only after confirming no process uses `skinbasenova-database-` (`HORIZON_PREFIX`, `REDIS_PREFIX` on all PHP/Horizon). Then `UNLINK` those keys. Do not delete `skinbase-database-queues:forum-moderation`.
---
## 10. Files changed
```text
app/Jobs/Concerns/RestoresAfterArtworkIdCursor.php
app/Jobs/RecComputeSimilarByTagsJob.php
app/Jobs/RecComputeSimilarByBehaviorJob.php
app/Jobs/RecComputeSimilarHybridJob.php
app/Jobs/IndexArtworkJob.php
app/Jobs/IndexUserJob.php
app/Jobs/RankBuildScopeListsJob.php
app/Services/Traffic/OnlineVisitorRepository.php
app/Services/CollectionBackgroundJobService.php
app/Console/Commands/DispatchCollectionMaintenanceCommand.php
routes/console.php
config/collections.php
config/forum_security.php
config/skinbase_ai_moderation.php
config/traffic.php
.env.example
phpunit.xml
tests/...
docs/optimization-m7.1-queue-compatibility.md
```
---
## 11. Tests
15 passed: RecCompute unserialize, cursor batching, Index* search queue, ranking uniqueId per scope, presence SSCAN/no SMEMBERS, collection dispatch gate.
---
## 12. Deploy sequence (later; not now)
1. Ship this release (jobs, uniqueness, presence SSCAN, collection gate, purge command).
2. Production `.env`:
```text
COLLECTIONS_V5_DISPATCH_ENABLED=false
RANK_SCOPE_JOB_UNIQUE_FOR=21600
```
Do not rely on `FORUM_QUEUE_DISPATCH_ENABLED` until the Forum plugin is changed.
3. `php artisan config:cache`
4. `sudo supervisorctl restart skinbase-horizon`
5. Optional: `sudo systemctl reload php8.4-fpm`
6. **After Horizon is on the new code**, purge stale Tags (operator, not cron):
```bash
php artisan skinbase:purge-queued-rec-tags --queue=default --dry-run
php artisan skinbase:purge-queued-rec-tags --queue=default --execute
```
7. Confirm `queues:default` RecComputeSimilarByTagsJob count is 0. Leave ranking/index jobs.
8. Tonight’s scheduler may dispatch **one** Tags rebuild.
**Horizon restart: YES.** Do not retry `failed_jobs`. Do not `SMEMBERS` presence index. Do not `DEL` forum-moderation.
---
## 13. Post-deploy verification
```bash
php artisan tinker --execute="echo json_encode([
'default' => Redis::llen('queues:default'),
'search' => Redis::llen('queues:search'),
'collections' => Redis::llen('queues:collections'),
'forum' => Redis::llen('queues:forum-moderation'),
]);"
# After some traffic, search LLEN should rise (Index* jobs), default ranking unique should stop stacking
# Do not retry failed_jobs
# Do not SMEMBERS skinbase:presence:online:index
```
---
## 14. Risks
- 169 in-queue tag jobs become “start from id 0” batches; duplicate compute, extra load, no crash.
- Ranking unique skips a scope for up to 360s if a lock is stuck.
- Presence UI shows ≤2000 members until prune.
- Forum enqueue may continue if the plugin ignores `FORUM_QUEUE_DISPATCH_ENABLED`.
- Collection tests set dispatch enabled in phpunit; production default is false.