Ship production optimization M1-M12.5A: queues, metrics, HTTP observability, and vector search reliability.

Keep similar-ai from tripping the global circuit on a lone URL 502, clamp Qdrant search to 100, and add Server-Timing plus slow-request logging. Studio shared props, Academy S3 exists caching, heat chunking, and Redis/scheduler hygiene stay in this rollout.
This commit is contained in:
2026-08-25 07:58:47 +02:00
parent f52879edbb
commit 8a80aae21e
114 changed files with 9293 additions and 284 deletions
@@ -0,0 +1,290 @@
# M10 — Database Query & Slow Query Optimization
```text
STATUS: COMPLETE (read-only production audit + two local application fixes)
PRODUCTION: no ALTER / DROP / OPTIMIZE / MySQL config / restart
WHEN: 2026-08-23
```
Production investigation used `ssh server3` as Skinbase (`skinbase@localhost`). Slow log file and `performance_schema.events_statements_summary_by_digest` were **not readable** with this grant. Cost ranking is from `SHOW GLOBAL STATUS`, `SHOW TABLE STATUS`, `SHOW INDEX`, `EXPLAIN`, and Laravel source mapping.
---
## 1. Production DB health snapshot
| Metric | Value |
| ------ | ----- |
| Version | Percona Server 8.4.11-11 |
| Uptime | ~122233 s (~34 h) |
| Threads_connected / Threads_running | 12 / 2 |
| max_connections / Max_used_connections | 120 / 26 |
| Slow_queries (cumulative) | 14301 (`long_query_time=0.5`, `min_examined_row_limit=100`) |
| Created_tmp_disk_tables | 21 (low vs memory temps) |
| Select_full_join | 30772 |
| Select_scan | ~1.15M |
| Sort_merge_passes | 107022 |
| InnoDB buffer pool | 10 GB; hit ratio ~99.9996%; pages_data ~24% |
| Innodb_row_lock_waits | 40794; avg ~4 ms; max ~449 ms |
Buffer pool and connection headroom are healthy. Cumulative `Slow_queries` is **not** treated as proof of a current problem (uptime + historical counter).
---
## 2. Slow-log configuration
| Variable | Production |
| -------- | ---------- |
| slow_query_log | ON |
| slow_query_log_file | `/var/log/mysql/slow.log` |
| long_query_time | 0.5 |
| min_examined_row_limit | 100 |
| log_queries_not_using_indexes | (not changed; not used as a tuning lever) |
**Do not change** these during M10. The log file is **not readable** as the app user. `pt-query-digest` was not used.
---
## 3–4. Slow-log / Performance Schema top fingerprints
**Unavailable** (`SELECT` denied on `events_statements_summary_by_digest`; slow.log unreadable).
Substitutes: `EXPLAIN` on known high-volume SQL + code-path frequency.
---
## 5–7. Top queries by total latency / average / rows examined
Ranked by **evidence from EXPLAIN + schedule frequency**, not P_S.
| Fingerprint | Feature | Path | rows examined (EXPLAIN) | Index | Severity |
| ----------- | ------- | ---- | ----------------------- | ----- | -------- |
| DISTINCT `artwork_id` WHERE `bucket_hour` BETWEEN ? AND ? | Rising heat | scheduled `nova:recalculate-heat` | ~2.46M covering `idx_bucket_artwork` | Using index + temp | HIGH (was hydrating all rows in PHP) |
| SELECT snapshots WHERE `bucket_hour` BETWEEN AND `artwork_id` IN (all IDs) | Rising heat | same command (pre-fix) | ~2.46M | `idx_bucket_artwork` | HIGH — **fixed locally** (chunked) |
| GROUP BY `artwork_id` 30d window | monthly leaderboards | `LeaderboardService` / `ArtworkHourlySnapshotWindow` | ~9.12M type=index `uq_artwork_bucket` | unique | MEDIUM (grows with 30d table) |
| `artwork_tag` GROUP BY `tag_id` COUNT | rec tags IDF | `RecComputeSimilarByTagsJob::handle` every batch | ~112k covering | `artwork_tag_tag_id_index` | MEDIUM — **fixed locally** (cache) |
| public list `is_public`+`is_approved` ORDER BY | HTTP browse | artwork listing | uses `idx_artworks_browse` backward | OK | NO CHANGE |
| eligible snapshot IDs | hourly snapshot | `MetricsSnapshotHourlyCommand` | `artworks_is_approved_index` + stats PK | OK | NO CHANGE |
| forum `reviewed=false` LIMIT 250 | `forum:scan-posts` | PRIMARY range | 13k-scale table | OK | NO CHANGE |
| collections lifecycle WHERE dates | `collections:sync-lifecycle` every 10m | small tables | OK | LOW / NO CHANGE |
---
## 8. Mapping queries → Laravel
| Query | Source |
| ----- | ------ |
| Heat DISTINCT + snapshot load + `artwork_stats` upsert | `app/Console/Commands/RecalculateHeatCommand.php::handle` |
| Hourly eligible IDs + upsert | `app/Console/Commands/MetricsSnapshotHourlyCommand` |
| 30d window MAX/MIN deltas | `app/Services/Metrics/ArtworkHourlySnapshotWindow.php` used by `StudioMetricsService`, `CreatorStudioOverviewService`, `LeaderboardService`, `CreatorJourneyService` |
| Rank lists LIMIT 200 | `app/Jobs/RankBuildScopeListsJob::fetchCandidates` (`rank_artwork_scores`) |
| Tag IDF GROUP BY | `app/Jobs/RecComputeSimilarByTagsJob::tagFrequencies` |
| Search document | `app/Jobs/IndexArtworkJob::handle` (`with` user, group, tags, categories.contentType, stats, awardStat) |
| Forum scan | `packages/klevze/Plugins/Forum` + schedule `forum:scan-posts --limit=250` |
| Collections sync | `app/Console/Commands/SyncCollectionLifecycleCommand.php` |
---
## 9. N+1 findings
| Area | Finding | Action |
| ---- | ------- | ------ |
| IndexArtworkJob | Already eager-loads relations used by `toSearchableArray()` | NO CHANGE |
| IndexUserJob | Review: not a heat source; search queue empty | NO CHANGE |
| Heat command | Was one giant hydrate, not N+1 | chunked load |
| Studio / Journey | Window helper issues bounded per-user SQL | NO CHANGE |
| Public artwork list | `idx_artworks_browse`; no new eager-load without evidence | NO CHANGE |
| Academy / forum cards | No production digest proving N+1 | NO CHANGE (do not globally eager-load) |
---
## 10. Artwork list / detail
Public browse uses `idx_artworks_browse` (EXPLAIN backward index scan). Category/tag pivots have supporting indexes from batch1. **NO CHANGE** to list indexes.
---
## 11. Rankings / leaderboards
`RankBuildScopeListsJob` reads `rank_artwork_scores` with LIMIT 200 after M7.1 uniqueness. DB side is cheap relative to heat/snapshots. `leaderboards:refresh` 30d GROUP BY remains the expensive reader as the snapshot table grows — keep current SQL; do **not** FORCE INDEX yet. Do **not** redesign queue architecture.
---
## 12. Metric-snapshot findings
Table ~8.87M rows, ~1855 MB (data ~740 / index ~1115). Unique `(artwork_id, bucket_hour)`.
Indexes:
| Index | Columns | Verdict |
| ----- | ------- | ------- |
| PRIMARY | `id` | keep |
| `uq_artwork_bucket` UNIQUE | `artwork_id`, `bucket_hour` | keep — upsert + per-artwork window |
| `idx_artwork_bucket` | `artwork_id`, `bucket_hour` | **exact duplicate of unique** — DROP candidate |
| `idx_bucket_hour` | `bucket_hour` | left prefix of `idx_bucket_artwork` — **do not drop in this milestone** |
| `idx_bucket_artwork` | `bucket_hour`, `artwork_id` | keep — heat DISTINCT covering |
Do **not** reduce 30-day retention. Do **not** `OPTIMIZE TABLE`. Do **not** partition.
---
## 13. Forum / collections
- `forum:scan-posts --limit=250` stays **synchronous**. No new index (PRIMARY + limit). Do not re-enable forum queues.
- Collections maintenance **queue remains disabled**. `collections:sync-lifecycle` is cheap date scans. Every-10-minute sync is fine from DB.
---
## 14. Academy
No P_S digest pointing at Academy. No speculative indexes or global eager-load.
---
## 15. Search / recommendations
- `IndexArtworkJob` eager-load is sufficient (no internal N+1 of the listed relations).
- Tag IDF GROUP BY ran **once per batch job**; cached 1h locally.
- Do **not** create a rec snapshot table (M4).
---
## 16. Table / index inventory (production estimates)
| Table | Rows (approx) | Notes |
| ----- | ------------- | ----- |
| `artwork_metric_snapshots_hourly` | 8.87M | index > data; duplicate secondary |
| `artworks` | (browse index used) | `idx_artworks_browse` |
| `artwork_stats` | 1:1 artwork | PK `artwork_id`; heat upsert |
| `artwork_tag` | covering tag_id | IDF GROUP BY |
| forum posts | ~13k | scan LIMIT 250 |
| rec tables | unique artwork+type+version | M7.1 intact |
---
## 17. Proposed indexes
**None** for add. Equality+range combinations already have `uq_artwork_bucket` and `idx_bucket_artwork`.
---
## 18. Proposed redundant index removals
**DROP `idx_artwork_bucket`** only.
Migration: `database/migrations/2026_08_23_120000_drop_duplicate_artwork_metric_snapshot_index.php`
**Do not run on production in the application deploy.** Separate schema window.
Percona 8.4: secondary index drop is typically `ALGORITHM=INPLACE`, `LOCK=NONE`. Disk: reclaim is InnoDB-internal, not filesystem (same as M5 prune). Rollback: re-add the duplicate index (unnecessary).
`idx_bucket_hour`: **NO CHANGE** until a dedicated EXPLAIN proves no leftover readers after drop of the duplicate.
---
## 19. Write-amplification
~50k upserts/hour. Each extra secondary index updates on every write.
- Duplicate `idx_artwork_bucket`: **100% write cost, 0% extra read benefit**.
- Remaining secondaries: `idx_bucket_artwork` is justified by heat DISTINCT covering.
---
## 20. Pagination / COUNT
No production evidence of OFFSET thousands on artworks browse. Do **not** rewrite pagination. Exact Studio counts remain M5.1 semantics.
---
## 21. Lock / transaction findings
Row lock waits exist but average 4 ms. Heat upsert is keyed by `artwork_id` (not a single hot row). **No isolation-level change.** Heat no longer loads 2.4M rows into PHP, which reduces transaction/memory pressure around upserts.
---
## 22. MySQL config verdict
**NO CHANGE** to `innodb_buffer_pool_size`, `max_connections`, redo, tmp_table_size, table caches, optimizer. Pool not under pressure.
---
## 23. Ranked findings
| Severity | Item |
| -------- | ---- |
| HIGH | Heat command hydrating full 24h snapshot set — **fixed** (chunked SELECT + narrower columns) |
| MEDIUM | Tag IDF GROUP BY per rec batch — **fixed** (Cache::remember 1h) |
| MEDIUM | Duplicate `idx_artwork_bucket` write amp — **migration created, not executed on prod** |
| MEDIUM | 30d leaderboard GROUP BY ~9M rows growing to ~36M — monitor; no FORCE INDEX yet |
| LOW | Sort_merge_passes / Select_scan cumulative — relate to heat DISTINCT temp; re-check after heat deploy |
| NO CHANGE | FPM max_children=14, Horizon counts, Redis, 30d retention, partition, OPTIMIZE, MySQL knobs, forum queues, rec snapshot table |
---
## 24. Files changed
- `app/Console/Commands/RecalculateHeatCommand.php`
- `app/Jobs/RecComputeSimilarByTagsJob.php`
- `tests/Feature/Metrics/RecalculateHeatCommandTest.php`
- `tests/Feature/Recommendations/RecComputeSimilarJobsTest.php`
- `database/migrations/2026_08_23_120000_drop_duplicate_artwork_metric_snapshot_index.php`
- `docs/optimization-m10-database-query-optimization.md`
---
## 25. Migrations created
`2026_08_23_120000_drop_duplicate_artwork_metric_snapshot_index.php` — **local only until a separate production schema deploy**.
---
## 26. Tests / results
```text
pest tests/Feature/Metrics/RecalculateHeatCommandTest.php PASS
phpunit --filter tag_idf RecComputeSimilarJobsTest OK
```
---
## 27. Expected benefit
- Heat: peak PHP memory and rows hydrated drop from ~all 24h snapshots (~1–2M rows) to `--chunk` × hours (default 1000 × 25).
- Tag jobs: one `artwork_tag` GROUP BY per hour per model version instead of once per batch/chain hop.
- Duplicate index drop (later): less write/storage on 50k rows/hour; index size currently already > data.
---
## 28. Application deployment sequence
1. Deploy code (heat chunking + IDF cache). **Do not** run the DROP INDEX migration.
2. After one `nova:recalculate-heat` cycle: confirm command duration and RSS vs previous (scheduler runtime).
3. After rec tags chain: confirm `artwork_tag` GROUP BY frequency (slow log / P_S if grants added later).
4. FPM/Horizon/Redis unchanged.
---
## 29. Schema deployment sequence (separate)
1. Confirm `SHOW INDEX FROM artwork_metric_snapshots_hourly` still shows `uq_artwork_bucket` and `idx_artwork_bucket` identical columns.
2. `ALTER TABLE ... DROP INDEX idx_artwork_bucket` via the migration in a dedicated window (INPLACE, no app coupling).
3. Re-EXPLAIN heat DISTINCT — expect `idx_bucket_artwork` unchanged.
4. Do not drop `idx_bucket_hour` in the same change.
---
## 30. Rollback
- App: revert heat/IDF commits; heat semantics unchanged (same formula, same upsert keys).
- Cache: `Cache::forget('rec:tag-idf:'.$modelVersion)` or wait TTL.
- Index: re-create `idx_artwork_bucket` only if a reader somehow used the name (none found; unique covers the same left prefix).
---
## Production verification plan (per fix)
| Fix | Verify |
| --- | ------ |
| Heat chunking | scheduler runtime; rows examined still on covering index; PHP memory; `artwork_stats.heat_score` still updates |
| IDF cache | query count of `artwork_tag` GROUP BY across a tags job chain |
| DROP duplicate index (later) | EXPLAIN heat + Studio window + prune `WHERE bucket_hour <` before/after |