Ship production optimization M1-M12.5A: queues, metrics, HTTP observability, and vector search reliability.

Keep similar-ai from tripping the global circuit on a lone URL 502, clamp Qdrant search to 100, and add Server-Timing plus slow-request logging. Studio shared props, Academy S3 exists caching, heat chunking, and Redis/scheduler hygiene stay in this rollout.
This commit is contained in:
2026-08-25 07:58:47 +02:00
parent f52879edbb
commit 8a80aae21e
114 changed files with 9293 additions and 284 deletions
@@ -0,0 +1,178 @@
# M4 — Recommendation Snapshot & Read Path
**Snapshot table needed: NO**
Investigation only plus a regression test. No production changes. No new tables.
---
## 1. What “snapshot table” meant
Two different things were previously named “snapshot”:
| Table | Size (prod) | Role |
| ----- | ----------: | ---- |
| `artwork_metric_snapshots_hourly` | **1855 MB / 8.6M rows** | Heat/ranking hourly totals (M0 DB residual). **Not** similar-art serving. |
| `rec_artwork_recs` | **11.6 MB / 27.7k rows** | Precomputed similar lists (JSON `recs`) read by the API. |
M4 is about the **recommendation read path**, not heat snapshots. A second copy of `rec_artwork_recs` would not address the 1.85 GB hourly table. That remains a later ranking/ops milestone.
---
## 2. Current data model
```text
rec_item_pairs
(a_artwork_id, b_artwork_id) UNIQUE
weight, updated_at
~74k rows, 16 MB
Written by RecBuildItemPairsFromFavouritesJob (DELETE all, then upsert chunks)
rec_artwork_recs
UNIQUE (artwork_id, rec_type, model_version)
recs JSON (ordered IDs), computed_at
rec_type: similar_tags | similar_behavior | similar_hybrid | similar_visual
Written by RecComputeSimilar* via updateOrCreate (one row at a time)
user_recommendation_cache
personalized “For You” blobs (~67 rows, 2.5 MB) — separate from similar-art
```
Frontend/API similar-art:
```text
GET /api/art/{id}/similar
GET /art/{id}/similar
→ HybridSimilarArtworksService::forArtwork
Cache::remember rec:artwork:{id}:similar:{model} TTL 6h
SELECT rec_artwork_recs WHERE artwork_id + rec_type + model_version (const)
hydrate Artwork whereIn(ids) public+published
```
---
## 3. Production row counts (2026-08-23, read-only)
| rec_type | rows | computed_at span |
| -------- | ---: | ---------------- |
| similar_hybrid | 13902 | 2026-03-15 → 2026-08-23 14:36 |
| similar_tags | 9479 | 2026-03-15 → 2026-08-23 14:58 |
| similar_behavior | 4851 | 2026-04-21 → 2026-08-23 14:34 |
~50k public artworks exist; lists are **incomplete** because nightly catalog jobs were dying (M1), not because of missing snapshots. Coverage will grow after M1 deploy.
---
## 4. Read/write flow
**Read:** unique const lookup (1 row) + Redis/cache 6h + one `whereIn` hydrate. Cheap.
**Write (M1 batches):** 200 artworks/job, `updateOrCreate` per artwork. No truncate of `rec_artwork_recs`. Unique key replaces one JSON list atomically.
**Write (pairs):** `DELETE FROM rec_item_pairs` then rebuild every 4 hours. Readers of pairs are **rebuild jobs**, not the HTTP similar endpoint.
**Partial visibility:** during a catalog rebuild, some artworks have today’s `computed_at`, others yesterday. HTTP cache may keep old IDs for 6h. There is no empty-table window for `rec_artwork_recs`. Mixing generations is **per artwork**, not a torn JSON array.
**Transactions:** none wrapping the full catalog. Each `updateOrCreate` is its own statement.
---
## 5. Query plans (production EXPLAIN)
Read:
```text
type=const
key=rec_artwork_recs_artwork_id_rec_type_model_version_unique
rows=1
```
Pairs UNION used by behavior rebuild (not HTTP):
```text
ref on unique (a_artwork_id, b_artwork_id)
ref on b_artwork_id index
UNION RESULT: Using temporary; Using filesort (tiny LIMIT 90)
```
No full table scan on the similar-art HTTP path.
---
## 6. Bottleneck (exact)
| Concern | Present? |
| ------- | -------- |
| Expensive HTTP reads | **No** — const unique + cache |
| Partial rec JSON during rebuild | **No** — row replace |
| Mixed old/new *across* artworks during overnight batch | **Yes, benign** with 6h cache |
| delete/reinsert on `rec_artwork_recs` | **No** |
| delete/reinsert on `rec_item_pairs` | **Yes**, 74k rows / 4h; not on the HTTP path |
| Poor indexing | **No** on read |
| Excessive Redis | One key per artwork, 6h TTL — normal |
| Dual generations mixed in one response | **No** |
**Exact bottleneck for similar-art quality:** incomplete `rec_artwork_recs` coverage from failed nightly jobs (M1), not read-path schema.
**Separate large table:** `artwork_metric_snapshots_hourly` 1.85 GB — ranking/heat, **out of M4 schema scope**.
---
## 7. Snapshot table?
```text
YES/NO: NO
```
Would double a **12 MB** table, add cutover complexity, and not fix M1 coverage or the 1.85 GB heat snapshots.
### Simpler alternatives (do not implement in M4 unless needed later)
1. **Deploy M1** so batches complete; coverage grows via `updateOrCreate`.
2. If `rec_item_pairs` DELETE causes empty-pair windows during overlap with behavior jobs: replace global delete with rebuild-into-temp + `RENAME TABLE` **on that 16 MB table only** — still not a rec snapshot table.
3. Heat snapshots: partition/prune `artwork_metric_snapshots_hourly` in a dedicated milestone.
---
## 8. Proposed schema if YES
N/A.
---
## 9–12. Migration / cutover / cleanup / deploy
N/A for snapshot tables. M1 deploy remains the rec-job fix.
Rollback: none (no schema).
---
## 13. Tests
Added: `keeps other rec types and other artworks visible when one row is rebuilt` in `tests/Feature/Recommendations/SimilarArtworksHybridTest.php`.
Existing hybrid tests already cover fallback chain, order, author cap, API `type=`.
---
## 14. Files changed
```text
tests/Feature/Recommendations/SimilarArtworksHybridTest.php
docs/optimization-m4-recommendation-snapshots.md
```
No migrations, no rec model changes.
---
## 15. Risks
None from schema (none added). Residual: pairs `DELETE` vs behavior schedule overlap; 6h cache lag after rebuild; heat snapshot table still large.
---
## Production
Read-only EXPLAIN + `information_schema` + grouped counts. Temp `/tmp` script removed. No deploys, no writes to rec tables.