Files
SkinbaseNova/docs/optimization-m4-recommendation-snapshots.md
T
klevze 8a80aae21e Ship production optimization M1-M12.5A: queues, metrics, HTTP observability, and vector search reliability.
Keep similar-ai from tripping the global circuit on a lone URL 502, clamp Qdrant search to 100, and add Server-Timing plus slow-request logging. Studio shared props, Academy S3 exists caching, heat chunking, and Redis/scheduler hygiene stay in this rollout.
2026-08-25 07:58:47 +02:00

179 lines
5.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# M4 — Recommendation Snapshot & Read Path
**Snapshot table needed: NO**
Investigation only plus a regression test. No production changes. No new tables.
---
## 1. What “snapshot table” meant
Two different things were previously named “snapshot”:
| Table | Size (prod) | Role |
| ----- | ----------: | ---- |
| `artwork_metric_snapshots_hourly` | **1855 MB / 8.6M rows** | Heat/ranking hourly totals (M0 DB residual). **Not** similar-art serving. |
| `rec_artwork_recs` | **11.6 MB / 27.7k rows** | Precomputed similar lists (JSON `recs`) read by the API. |
M4 is about the **recommendation read path**, not heat snapshots. A second copy of `rec_artwork_recs` would not address the 1.85 GB hourly table. That remains a later ranking/ops milestone.
---
## 2. Current data model
```text
rec_item_pairs
(a_artwork_id, b_artwork_id) UNIQUE
weight, updated_at
~74k rows, 16 MB
Written by RecBuildItemPairsFromFavouritesJob (DELETE all, then upsert chunks)
rec_artwork_recs
UNIQUE (artwork_id, rec_type, model_version)
recs JSON (ordered IDs), computed_at
rec_type: similar_tags | similar_behavior | similar_hybrid | similar_visual
Written by RecComputeSimilar* via updateOrCreate (one row at a time)
user_recommendation_cache
personalized “For You” blobs (~67 rows, 2.5 MB) — separate from similar-art
```
Frontend/API similar-art:
```text
GET /api/art/{id}/similar
GET /art/{id}/similar
→ HybridSimilarArtworksService::forArtwork
Cache::remember rec:artwork:{id}:similar:{model} TTL 6h
SELECT rec_artwork_recs WHERE artwork_id + rec_type + model_version (const)
hydrate Artwork whereIn(ids) public+published
```
---
## 3. Production row counts (2026-08-23, read-only)
| rec_type | rows | computed_at span |
| -------- | ---: | ---------------- |
| similar_hybrid | 13902 | 2026-03-15 → 2026-08-23 14:36 |
| similar_tags | 9479 | 2026-03-15 → 2026-08-23 14:58 |
| similar_behavior | 4851 | 2026-04-21 → 2026-08-23 14:34 |
~50k public artworks exist; lists are **incomplete** because nightly catalog jobs were dying (M1), not because of missing snapshots. Coverage will grow after M1 deploy.
---
## 4. Read/write flow
**Read:** unique const lookup (1 row) + Redis/cache 6h + one `whereIn` hydrate. Cheap.
**Write (M1 batches):** 200 artworks/job, `updateOrCreate` per artwork. No truncate of `rec_artwork_recs`. Unique key replaces one JSON list atomically.
**Write (pairs):** `DELETE FROM rec_item_pairs` then rebuild every 4 hours. Readers of pairs are **rebuild jobs**, not the HTTP similar endpoint.
**Partial visibility:** during a catalog rebuild, some artworks have today’s `computed_at`, others yesterday. HTTP cache may keep old IDs for 6h. There is no empty-table window for `rec_artwork_recs`. Mixing generations is **per artwork**, not a torn JSON array.
**Transactions:** none wrapping the full catalog. Each `updateOrCreate` is its own statement.
---
## 5. Query plans (production EXPLAIN)
Read:
```text
type=const
key=rec_artwork_recs_artwork_id_rec_type_model_version_unique
rows=1
```
Pairs UNION used by behavior rebuild (not HTTP):
```text
ref on unique (a_artwork_id, b_artwork_id)
ref on b_artwork_id index
UNION RESULT: Using temporary; Using filesort (tiny LIMIT 90)
```
No full table scan on the similar-art HTTP path.
---
## 6. Bottleneck (exact)
| Concern | Present? |
| ------- | -------- |
| Expensive HTTP reads | **No** — const unique + cache |
| Partial rec JSON during rebuild | **No** — row replace |
| Mixed old/new *across* artworks during overnight batch | **Yes, benign** with 6h cache |
| delete/reinsert on `rec_artwork_recs` | **No** |
| delete/reinsert on `rec_item_pairs` | **Yes**, 74k rows / 4h; not on the HTTP path |
| Poor indexing | **No** on read |
| Excessive Redis | One key per artwork, 6h TTL — normal |
| Dual generations mixed in one response | **No** |
**Exact bottleneck for similar-art quality:** incomplete `rec_artwork_recs` coverage from failed nightly jobs (M1), not read-path schema.
**Separate large table:** `artwork_metric_snapshots_hourly` 1.85 GB — ranking/heat, **out of M4 schema scope**.
---
## 7. Snapshot table?
```text
YES/NO: NO
```
Would double a **12 MB** table, add cutover complexity, and not fix M1 coverage or the 1.85 GB heat snapshots.
### Simpler alternatives (do not implement in M4 unless needed later)
1. **Deploy M1** so batches complete; coverage grows via `updateOrCreate`.
2. If `rec_item_pairs` DELETE causes empty-pair windows during overlap with behavior jobs: replace global delete with rebuild-into-temp + `RENAME TABLE` **on that 16 MB table only** — still not a rec snapshot table.
3. Heat snapshots: partition/prune `artwork_metric_snapshots_hourly` in a dedicated milestone.
---
## 8. Proposed schema if YES
N/A.
---
## 9–12. Migration / cutover / cleanup / deploy
N/A for snapshot tables. M1 deploy remains the rec-job fix.
Rollback: none (no schema).
---
## 13. Tests
Added: `keeps other rec types and other artworks visible when one row is rebuilt` in `tests/Feature/Recommendations/SimilarArtworksHybridTest.php`.
Existing hybrid tests already cover fallback chain, order, author cap, API `type=`.
---
## 14. Files changed
```text
tests/Feature/Recommendations/SimilarArtworksHybridTest.php
docs/optimization-m4-recommendation-snapshots.md
```
No migrations, no rec model changes.
---
## 15. Risks
None from schema (none added). Residual: pairs `DELETE` vs behavior schedule overlap; 6h cache lag after rebuild; heat snapshot table still large.
---
## Production
Read-only EXPLAIN + `information_schema` + grouped counts. Temp `/tmp` script removed. No deploys, no writes to rec tables.