Files
SkinbaseNova/docs/optimization-m10-database-query-optimization.md
T
klevze 8a80aae21e Ship production optimization M1-M12.5A: queues, metrics, HTTP observability, and vector search reliability.
Keep similar-ai from tripping the global circuit on a lone URL 502, clamp Qdrant search to 100, and add Server-Timing plus slow-request logging. Studio shared props, Academy S3 exists caching, heat chunking, and Redis/scheduler hygiene stay in this rollout.
2026-08-25 07:58:47 +02:00

291 lines
12 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# M10 — Database Query & Slow Query Optimization
```text
STATUS: COMPLETE (read-only production audit + two local application fixes)
PRODUCTION: no ALTER / DROP / OPTIMIZE / MySQL config / restart
WHEN: 2026-08-23
```
Production investigation used `ssh server3` as Skinbase (`skinbase@localhost`). Slow log file and `performance_schema.events_statements_summary_by_digest` were **not readable** with this grant. Cost ranking is from `SHOW GLOBAL STATUS`, `SHOW TABLE STATUS`, `SHOW INDEX`, `EXPLAIN`, and Laravel source mapping.
---
## 1. Production DB health snapshot
| Metric | Value |
| ------ | ----- |
| Version | Percona Server 8.4.11-11 |
| Uptime | ~122233 s (~34 h) |
| Threads_connected / Threads_running | 12 / 2 |
| max_connections / Max_used_connections | 120 / 26 |
| Slow_queries (cumulative) | 14301 (`long_query_time=0.5`, `min_examined_row_limit=100`) |
| Created_tmp_disk_tables | 21 (low vs memory temps) |
| Select_full_join | 30772 |
| Select_scan | ~1.15M |
| Sort_merge_passes | 107022 |
| InnoDB buffer pool | 10 GB; hit ratio ~99.9996%; pages_data ~24% |
| Innodb_row_lock_waits | 40794; avg ~4 ms; max ~449 ms |
Buffer pool and connection headroom are healthy. Cumulative `Slow_queries` is **not** treated as proof of a current problem (uptime + historical counter).
---
## 2. Slow-log configuration
| Variable | Production |
| -------- | ---------- |
| slow_query_log | ON |
| slow_query_log_file | `/var/log/mysql/slow.log` |
| long_query_time | 0.5 |
| min_examined_row_limit | 100 |
| log_queries_not_using_indexes | (not changed; not used as a tuning lever) |
**Do not change** these during M10. The log file is **not readable** as the app user. `pt-query-digest` was not used.
---
## 3–4. Slow-log / Performance Schema top fingerprints
**Unavailable** (`SELECT` denied on `events_statements_summary_by_digest`; slow.log unreadable).
Substitutes: `EXPLAIN` on known high-volume SQL + code-path frequency.
---
## 5–7. Top queries by total latency / average / rows examined
Ranked by **evidence from EXPLAIN + schedule frequency**, not P_S.
| Fingerprint | Feature | Path | rows examined (EXPLAIN) | Index | Severity |
| ----------- | ------- | ---- | ----------------------- | ----- | -------- |
| DISTINCT `artwork_id` WHERE `bucket_hour` BETWEEN ? AND ? | Rising heat | scheduled `nova:recalculate-heat` | ~2.46M covering `idx_bucket_artwork` | Using index + temp | HIGH (was hydrating all rows in PHP) |
| SELECT snapshots WHERE `bucket_hour` BETWEEN AND `artwork_id` IN (all IDs) | Rising heat | same command (pre-fix) | ~2.46M | `idx_bucket_artwork` | HIGH — **fixed locally** (chunked) |
| GROUP BY `artwork_id` 30d window | monthly leaderboards | `LeaderboardService` / `ArtworkHourlySnapshotWindow` | ~9.12M type=index `uq_artwork_bucket` | unique | MEDIUM (grows with 30d table) |
| `artwork_tag` GROUP BY `tag_id` COUNT | rec tags IDF | `RecComputeSimilarByTagsJob::handle` every batch | ~112k covering | `artwork_tag_tag_id_index` | MEDIUM — **fixed locally** (cache) |
| public list `is_public`+`is_approved` ORDER BY | HTTP browse | artwork listing | uses `idx_artworks_browse` backward | OK | NO CHANGE |
| eligible snapshot IDs | hourly snapshot | `MetricsSnapshotHourlyCommand` | `artworks_is_approved_index` + stats PK | OK | NO CHANGE |
| forum `reviewed=false` LIMIT 250 | `forum:scan-posts` | PRIMARY range | 13k-scale table | OK | NO CHANGE |
| collections lifecycle WHERE dates | `collections:sync-lifecycle` every 10m | small tables | OK | LOW / NO CHANGE |
---
## 8. Mapping queries → Laravel
| Query | Source |
| ----- | ------ |
| Heat DISTINCT + snapshot load + `artwork_stats` upsert | `app/Console/Commands/RecalculateHeatCommand.php::handle` |
| Hourly eligible IDs + upsert | `app/Console/Commands/MetricsSnapshotHourlyCommand` |
| 30d window MAX/MIN deltas | `app/Services/Metrics/ArtworkHourlySnapshotWindow.php` used by `StudioMetricsService`, `CreatorStudioOverviewService`, `LeaderboardService`, `CreatorJourneyService` |
| Rank lists LIMIT 200 | `app/Jobs/RankBuildScopeListsJob::fetchCandidates` (`rank_artwork_scores`) |
| Tag IDF GROUP BY | `app/Jobs/RecComputeSimilarByTagsJob::tagFrequencies` |
| Search document | `app/Jobs/IndexArtworkJob::handle` (`with` user, group, tags, categories.contentType, stats, awardStat) |
| Forum scan | `packages/klevze/Plugins/Forum` + schedule `forum:scan-posts --limit=250` |
| Collections sync | `app/Console/Commands/SyncCollectionLifecycleCommand.php` |
---
## 9. N+1 findings
| Area | Finding | Action |
| ---- | ------- | ------ |
| IndexArtworkJob | Already eager-loads relations used by `toSearchableArray()` | NO CHANGE |
| IndexUserJob | Review: not a heat source; search queue empty | NO CHANGE |
| Heat command | Was one giant hydrate, not N+1 | chunked load |
| Studio / Journey | Window helper issues bounded per-user SQL | NO CHANGE |
| Public artwork list | `idx_artworks_browse`; no new eager-load without evidence | NO CHANGE |
| Academy / forum cards | No production digest proving N+1 | NO CHANGE (do not globally eager-load) |
---
## 10. Artwork list / detail
Public browse uses `idx_artworks_browse` (EXPLAIN backward index scan). Category/tag pivots have supporting indexes from batch1. **NO CHANGE** to list indexes.
---
## 11. Rankings / leaderboards
`RankBuildScopeListsJob` reads `rank_artwork_scores` with LIMIT 200 after M7.1 uniqueness. DB side is cheap relative to heat/snapshots. `leaderboards:refresh` 30d GROUP BY remains the expensive reader as the snapshot table grows — keep current SQL; do **not** FORCE INDEX yet. Do **not** redesign queue architecture.
---
## 12. Metric-snapshot findings
Table ~8.87M rows, ~1855 MB (data ~740 / index ~1115). Unique `(artwork_id, bucket_hour)`.
Indexes:
| Index | Columns | Verdict |
| ----- | ------- | ------- |
| PRIMARY | `id` | keep |
| `uq_artwork_bucket` UNIQUE | `artwork_id`, `bucket_hour` | keep — upsert + per-artwork window |
| `idx_artwork_bucket` | `artwork_id`, `bucket_hour` | **exact duplicate of unique** — DROP candidate |
| `idx_bucket_hour` | `bucket_hour` | left prefix of `idx_bucket_artwork` — **do not drop in this milestone** |
| `idx_bucket_artwork` | `bucket_hour`, `artwork_id` | keep — heat DISTINCT covering |
Do **not** reduce 30-day retention. Do **not** `OPTIMIZE TABLE`. Do **not** partition.
---
## 13. Forum / collections
- `forum:scan-posts --limit=250` stays **synchronous**. No new index (PRIMARY + limit). Do not re-enable forum queues.
- Collections maintenance **queue remains disabled**. `collections:sync-lifecycle` is cheap date scans. Every-10-minute sync is fine from DB.
---
## 14. Academy
No P_S digest pointing at Academy. No speculative indexes or global eager-load.
---
## 15. Search / recommendations
- `IndexArtworkJob` eager-load is sufficient (no internal N+1 of the listed relations).
- Tag IDF GROUP BY ran **once per batch job**; cached 1h locally.
- Do **not** create a rec snapshot table (M4).
---
## 16. Table / index inventory (production estimates)
| Table | Rows (approx) | Notes |
| ----- | ------------- | ----- |
| `artwork_metric_snapshots_hourly` | 8.87M | index > data; duplicate secondary |
| `artworks` | (browse index used) | `idx_artworks_browse` |
| `artwork_stats` | 1:1 artwork | PK `artwork_id`; heat upsert |
| `artwork_tag` | covering tag_id | IDF GROUP BY |
| forum posts | ~13k | scan LIMIT 250 |
| rec tables | unique artwork+type+version | M7.1 intact |
---
## 17. Proposed indexes
**None** for add. Equality+range combinations already have `uq_artwork_bucket` and `idx_bucket_artwork`.
---
## 18. Proposed redundant index removals
**DROP `idx_artwork_bucket`** only.
Migration: `database/migrations/2026_08_23_120000_drop_duplicate_artwork_metric_snapshot_index.php`
**Do not run on production in the application deploy.** Separate schema window.
Percona 8.4: secondary index drop is typically `ALGORITHM=INPLACE`, `LOCK=NONE`. Disk: reclaim is InnoDB-internal, not filesystem (same as M5 prune). Rollback: re-add the duplicate index (unnecessary).
`idx_bucket_hour`: **NO CHANGE** until a dedicated EXPLAIN proves no leftover readers after drop of the duplicate.
---
## 19. Write-amplification
~50k upserts/hour. Each extra secondary index updates on every write.
- Duplicate `idx_artwork_bucket`: **100% write cost, 0% extra read benefit**.
- Remaining secondaries: `idx_bucket_artwork` is justified by heat DISTINCT covering.
---
## 20. Pagination / COUNT
No production evidence of OFFSET thousands on artworks browse. Do **not** rewrite pagination. Exact Studio counts remain M5.1 semantics.
---
## 21. Lock / transaction findings
Row lock waits exist but average 4 ms. Heat upsert is keyed by `artwork_id` (not a single hot row). **No isolation-level change.** Heat no longer loads 2.4M rows into PHP, which reduces transaction/memory pressure around upserts.
---
## 22. MySQL config verdict
**NO CHANGE** to `innodb_buffer_pool_size`, `max_connections`, redo, tmp_table_size, table caches, optimizer. Pool not under pressure.
---
## 23. Ranked findings
| Severity | Item |
| -------- | ---- |
| HIGH | Heat command hydrating full 24h snapshot set — **fixed** (chunked SELECT + narrower columns) |
| MEDIUM | Tag IDF GROUP BY per rec batch — **fixed** (Cache::remember 1h) |
| MEDIUM | Duplicate `idx_artwork_bucket` write amp — **migration created, not executed on prod** |
| MEDIUM | 30d leaderboard GROUP BY ~9M rows growing to ~36M — monitor; no FORCE INDEX yet |
| LOW | Sort_merge_passes / Select_scan cumulative — relate to heat DISTINCT temp; re-check after heat deploy |
| NO CHANGE | FPM max_children=14, Horizon counts, Redis, 30d retention, partition, OPTIMIZE, MySQL knobs, forum queues, rec snapshot table |
---
## 24. Files changed
- `app/Console/Commands/RecalculateHeatCommand.php`
- `app/Jobs/RecComputeSimilarByTagsJob.php`
- `tests/Feature/Metrics/RecalculateHeatCommandTest.php`
- `tests/Feature/Recommendations/RecComputeSimilarJobsTest.php`
- `database/migrations/2026_08_23_120000_drop_duplicate_artwork_metric_snapshot_index.php`
- `docs/optimization-m10-database-query-optimization.md`
---
## 25. Migrations created
`2026_08_23_120000_drop_duplicate_artwork_metric_snapshot_index.php` — **local only until a separate production schema deploy**.
---
## 26. Tests / results
```text
pest tests/Feature/Metrics/RecalculateHeatCommandTest.php PASS
phpunit --filter tag_idf RecComputeSimilarJobsTest OK
```
---
## 27. Expected benefit
- Heat: peak PHP memory and rows hydrated drop from ~all 24h snapshots (~1–2M rows) to `--chunk` × hours (default 1000 × 25).
- Tag jobs: one `artwork_tag` GROUP BY per hour per model version instead of once per batch/chain hop.
- Duplicate index drop (later): less write/storage on 50k rows/hour; index size currently already > data.
---
## 28. Application deployment sequence
1. Deploy code (heat chunking + IDF cache). **Do not** run the DROP INDEX migration.
2. After one `nova:recalculate-heat` cycle: confirm command duration and RSS vs previous (scheduler runtime).
3. After rec tags chain: confirm `artwork_tag` GROUP BY frequency (slow log / P_S if grants added later).
4. FPM/Horizon/Redis unchanged.
---
## 29. Schema deployment sequence (separate)
1. Confirm `SHOW INDEX FROM artwork_metric_snapshots_hourly` still shows `uq_artwork_bucket` and `idx_artwork_bucket` identical columns.
2. `ALTER TABLE ... DROP INDEX idx_artwork_bucket` via the migration in a dedicated window (INPLACE, no app coupling).
3. Re-EXPLAIN heat DISTINCT — expect `idx_bucket_artwork` unchanged.
4. Do not drop `idx_bucket_hour` in the same change.
---
## 30. Rollback
- App: revert heat/IDF commits; heat semantics unchanged (same formula, same upsert keys).
- Cache: `Cache::forget('rec:tag-idf:'.$modelVersion)` or wait TTL.
- Index: re-create `idx_artwork_bucket` only if a reader somehow used the name (none found; unique covers the same left prefix).
---
## Production verification plan (per fix)
| Fix | Verify |
| --- | ------ |
| Heat chunking | scheduler runtime; rows examined still on covering index; PHP memory; `artwork_stats.heat_score` still updates |
| IDF cache | query count of `artwork_tag` GROUP BY across a tags job chain |
| DROP duplicate index (later) | EXPLAIN heat + Studio window + prune `WHERE bucket_hour <` before/after |