Ship production optimization M1-M12.5A: queues, metrics, HTTP observability, and vector search reliability.
Keep similar-ai from tripping the global circuit on a lone URL 502, clamp Qdrant search to 100, and add Server-Timing plus slow-request logging. Studio shared props, Academy S3 exists caching, heat chunking, and Redis/scheduler hygiene stay in this rollout.
This commit is contained in:
@@ -0,0 +1,178 @@
|
||||
# M4 — Recommendation Snapshot & Read Path
|
||||
|
||||
**Snapshot table needed: NO**
|
||||
|
||||
Investigation only plus a regression test. No production changes. No new tables.
|
||||
|
||||
---
|
||||
|
||||
## 1. What “snapshot table” meant
|
||||
|
||||
Two different things were previously named “snapshot”:
|
||||
|
||||
| Table | Size (prod) | Role |
|
||||
| ----- | ----------: | ---- |
|
||||
| `artwork_metric_snapshots_hourly` | **1855 MB / 8.6M rows** | Heat/ranking hourly totals (M0 DB residual). **Not** similar-art serving. |
|
||||
| `rec_artwork_recs` | **11.6 MB / 27.7k rows** | Precomputed similar lists (JSON `recs`) read by the API. |
|
||||
|
||||
M4 is about the **recommendation read path**, not heat snapshots. A second copy of `rec_artwork_recs` would not address the 1.85 GB hourly table. That remains a later ranking/ops milestone.
|
||||
|
||||
---
|
||||
|
||||
## 2. Current data model
|
||||
|
||||
```text
|
||||
rec_item_pairs
|
||||
(a_artwork_id, b_artwork_id) UNIQUE
|
||||
weight, updated_at
|
||||
~74k rows, 16 MB
|
||||
Written by RecBuildItemPairsFromFavouritesJob (DELETE all, then upsert chunks)
|
||||
|
||||
rec_artwork_recs
|
||||
UNIQUE (artwork_id, rec_type, model_version)
|
||||
recs JSON (ordered IDs), computed_at
|
||||
rec_type: similar_tags | similar_behavior | similar_hybrid | similar_visual
|
||||
Written by RecComputeSimilar* via updateOrCreate (one row at a time)
|
||||
|
||||
user_recommendation_cache
|
||||
personalized “For You” blobs (~67 rows, 2.5 MB) — separate from similar-art
|
||||
```
|
||||
|
||||
Frontend/API similar-art:
|
||||
|
||||
```text
|
||||
GET /api/art/{id}/similar
|
||||
GET /art/{id}/similar
|
||||
→ HybridSimilarArtworksService::forArtwork
|
||||
Cache::remember rec:artwork:{id}:similar:{model} TTL 6h
|
||||
SELECT rec_artwork_recs WHERE artwork_id + rec_type + model_version (const)
|
||||
hydrate Artwork whereIn(ids) public+published
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 3. Production row counts (2026-08-23, read-only)
|
||||
|
||||
| rec_type | rows | computed_at span |
|
||||
| -------- | ---: | ---------------- |
|
||||
| similar_hybrid | 13902 | 2026-03-15 → 2026-08-23 14:36 |
|
||||
| similar_tags | 9479 | 2026-03-15 → 2026-08-23 14:58 |
|
||||
| similar_behavior | 4851 | 2026-04-21 → 2026-08-23 14:34 |
|
||||
|
||||
~50k public artworks exist; lists are **incomplete** because nightly catalog jobs were dying (M1), not because of missing snapshots. Coverage will grow after M1 deploy.
|
||||
|
||||
---
|
||||
|
||||
## 4. Read/write flow
|
||||
|
||||
**Read:** unique const lookup (1 row) + Redis/cache 6h + one `whereIn` hydrate. Cheap.
|
||||
|
||||
**Write (M1 batches):** 200 artworks/job, `updateOrCreate` per artwork. No truncate of `rec_artwork_recs`. Unique key replaces one JSON list atomically.
|
||||
|
||||
**Write (pairs):** `DELETE FROM rec_item_pairs` then rebuild every 4 hours. Readers of pairs are **rebuild jobs**, not the HTTP similar endpoint.
|
||||
|
||||
**Partial visibility:** during a catalog rebuild, some artworks have today’s `computed_at`, others yesterday. HTTP cache may keep old IDs for 6h. There is no empty-table window for `rec_artwork_recs`. Mixing generations is **per artwork**, not a torn JSON array.
|
||||
|
||||
**Transactions:** none wrapping the full catalog. Each `updateOrCreate` is its own statement.
|
||||
|
||||
---
|
||||
|
||||
## 5. Query plans (production EXPLAIN)
|
||||
|
||||
Read:
|
||||
|
||||
```text
|
||||
type=const
|
||||
key=rec_artwork_recs_artwork_id_rec_type_model_version_unique
|
||||
rows=1
|
||||
```
|
||||
|
||||
Pairs UNION used by behavior rebuild (not HTTP):
|
||||
|
||||
```text
|
||||
ref on unique (a_artwork_id, b_artwork_id)
|
||||
ref on b_artwork_id index
|
||||
UNION RESULT: Using temporary; Using filesort (tiny LIMIT 90)
|
||||
```
|
||||
|
||||
No full table scan on the similar-art HTTP path.
|
||||
|
||||
---
|
||||
|
||||
## 6. Bottleneck (exact)
|
||||
|
||||
| Concern | Present? |
|
||||
| ------- | -------- |
|
||||
| Expensive HTTP reads | **No** — const unique + cache |
|
||||
| Partial rec JSON during rebuild | **No** — row replace |
|
||||
| Mixed old/new *across* artworks during overnight batch | **Yes, benign** with 6h cache |
|
||||
| delete/reinsert on `rec_artwork_recs` | **No** |
|
||||
| delete/reinsert on `rec_item_pairs` | **Yes**, 74k rows / 4h; not on the HTTP path |
|
||||
| Poor indexing | **No** on read |
|
||||
| Excessive Redis | One key per artwork, 6h TTL — normal |
|
||||
| Dual generations mixed in one response | **No** |
|
||||
|
||||
**Exact bottleneck for similar-art quality:** incomplete `rec_artwork_recs` coverage from failed nightly jobs (M1), not read-path schema.
|
||||
|
||||
**Separate large table:** `artwork_metric_snapshots_hourly` 1.85 GB — ranking/heat, **out of M4 schema scope**.
|
||||
|
||||
---
|
||||
|
||||
## 7. Snapshot table?
|
||||
|
||||
```text
|
||||
YES/NO: NO
|
||||
```
|
||||
|
||||
Would double a **12 MB** table, add cutover complexity, and not fix M1 coverage or the 1.85 GB heat snapshots.
|
||||
|
||||
### Simpler alternatives (do not implement in M4 unless needed later)
|
||||
|
||||
1. **Deploy M1** so batches complete; coverage grows via `updateOrCreate`.
|
||||
2. If `rec_item_pairs` DELETE causes empty-pair windows during overlap with behavior jobs: replace global delete with rebuild-into-temp + `RENAME TABLE` **on that 16 MB table only** — still not a rec snapshot table.
|
||||
3. Heat snapshots: partition/prune `artwork_metric_snapshots_hourly` in a dedicated milestone.
|
||||
|
||||
---
|
||||
|
||||
## 8. Proposed schema if YES
|
||||
|
||||
N/A.
|
||||
|
||||
---
|
||||
|
||||
## 9–12. Migration / cutover / cleanup / deploy
|
||||
|
||||
N/A for snapshot tables. M1 deploy remains the rec-job fix.
|
||||
|
||||
Rollback: none (no schema).
|
||||
|
||||
---
|
||||
|
||||
## 13. Tests
|
||||
|
||||
Added: `keeps other rec types and other artworks visible when one row is rebuilt` in `tests/Feature/Recommendations/SimilarArtworksHybridTest.php`.
|
||||
|
||||
Existing hybrid tests already cover fallback chain, order, author cap, API `type=`.
|
||||
|
||||
---
|
||||
|
||||
## 14. Files changed
|
||||
|
||||
```text
|
||||
tests/Feature/Recommendations/SimilarArtworksHybridTest.php
|
||||
docs/optimization-m4-recommendation-snapshots.md
|
||||
```
|
||||
|
||||
No migrations, no rec model changes.
|
||||
|
||||
---
|
||||
|
||||
## 15. Risks
|
||||
|
||||
None from schema (none added). Residual: pairs `DELETE` vs behavior schedule overlap; 6h cache lag after rebuild; heat snapshot table still large.
|
||||
|
||||
---
|
||||
|
||||
## Production
|
||||
|
||||
Read-only EXPLAIN + `information_schema` + grouped counts. Temp `/tmp` script removed. No deploys, no writes to rec tables.
|
||||
Reference in New Issue
Block a user