Keep similar-ai from tripping the global circuit on a lone URL 502, clamp Qdrant search to 100, and add Server-Timing plus slow-request logging. Studio shared props, Academy S3 exists caching, heat chunking, and Redis/scheduler hygiene stay in this rollout.
12 KiB
M10 — Database Query & Slow Query Optimization
STATUS: COMPLETE (read-only production audit + two local application fixes)
PRODUCTION: no ALTER / DROP / OPTIMIZE / MySQL config / restart
WHEN: 2026-08-23
Production investigation used ssh server3 as Skinbase (skinbase@localhost). Slow log file and performance_schema.events_statements_summary_by_digest were not readable with this grant. Cost ranking is from SHOW GLOBAL STATUS, SHOW TABLE STATUS, SHOW INDEX, EXPLAIN, and Laravel source mapping.
1. Production DB health snapshot
| Metric | Value |
|---|---|
| Version | Percona Server 8.4.11-11 |
| Uptime | ~122233 s (~34 h) |
| Threads_connected / Threads_running | 12 / 2 |
| max_connections / Max_used_connections | 120 / 26 |
| Slow_queries (cumulative) | 14301 (long_query_time=0.5, min_examined_row_limit=100) |
| Created_tmp_disk_tables | 21 (low vs memory temps) |
| Select_full_join | 30772 |
| Select_scan | ~1.15M |
| Sort_merge_passes | 107022 |
| InnoDB buffer pool | 10 GB; hit ratio ~99.9996%; pages_data ~24% |
| Innodb_row_lock_waits | 40794; avg ~4 ms; max ~449 ms |
Buffer pool and connection headroom are healthy. Cumulative Slow_queries is not treated as proof of a current problem (uptime + historical counter).
2. Slow-log configuration
| Variable | Production |
|---|---|
| slow_query_log | ON |
| slow_query_log_file | /var/log/mysql/slow.log |
| long_query_time | 0.5 |
| min_examined_row_limit | 100 |
| log_queries_not_using_indexes | (not changed; not used as a tuning lever) |
Do not change these during M10. The log file is not readable as the app user. pt-query-digest was not used.
3–4. Slow-log / Performance Schema top fingerprints
Unavailable (SELECT denied on events_statements_summary_by_digest; slow.log unreadable).
Substitutes: EXPLAIN on known high-volume SQL + code-path frequency.
5–7. Top queries by total latency / average / rows examined
Ranked by evidence from EXPLAIN + schedule frequency, not P_S.
| Fingerprint | Feature | Path | rows examined (EXPLAIN) | Index | Severity |
|---|---|---|---|---|---|
DISTINCT artwork_id WHERE bucket_hour BETWEEN ? AND ? |
Rising heat | scheduled nova:recalculate-heat |
~2.46M covering idx_bucket_artwork |
Using index + temp | HIGH (was hydrating all rows in PHP) |
SELECT snapshots WHERE bucket_hour BETWEEN AND artwork_id IN (all IDs) |
Rising heat | same command (pre-fix) | ~2.46M | idx_bucket_artwork |
HIGH — fixed locally (chunked) |
GROUP BY artwork_id 30d window |
monthly leaderboards | LeaderboardService / ArtworkHourlySnapshotWindow |
~9.12M type=index uq_artwork_bucket |
unique | MEDIUM (grows with 30d table) |
artwork_tag GROUP BY tag_id COUNT |
rec tags IDF | RecComputeSimilarByTagsJob::handle every batch |
~112k covering | artwork_tag_tag_id_index |
MEDIUM — fixed locally (cache) |
public list is_public+is_approved ORDER BY |
HTTP browse | artwork listing | uses idx_artworks_browse backward |
OK | NO CHANGE |
| eligible snapshot IDs | hourly snapshot | MetricsSnapshotHourlyCommand |
artworks_is_approved_index + stats PK |
OK | NO CHANGE |
forum reviewed=false LIMIT 250 |
forum:scan-posts |
PRIMARY range | 13k-scale table | OK | NO CHANGE |
| collections lifecycle WHERE dates | collections:sync-lifecycle every 10m |
small tables | OK | LOW / NO CHANGE |
8. Mapping queries → Laravel
| Query | Source |
|---|---|
Heat DISTINCT + snapshot load + artwork_stats upsert |
app/Console/Commands/RecalculateHeatCommand.php::handle |
| Hourly eligible IDs + upsert | app/Console/Commands/MetricsSnapshotHourlyCommand |
| 30d window MAX/MIN deltas | app/Services/Metrics/ArtworkHourlySnapshotWindow.php used by StudioMetricsService, CreatorStudioOverviewService, LeaderboardService, CreatorJourneyService |
| Rank lists LIMIT 200 | app/Jobs/RankBuildScopeListsJob::fetchCandidates (rank_artwork_scores) |
| Tag IDF GROUP BY | app/Jobs/RecComputeSimilarByTagsJob::tagFrequencies |
| Search document | app/Jobs/IndexArtworkJob::handle (with user, group, tags, categories.contentType, stats, awardStat) |
| Forum scan | packages/klevze/Plugins/Forum + schedule forum:scan-posts --limit=250 |
| Collections sync | app/Console/Commands/SyncCollectionLifecycleCommand.php |
9. N+1 findings
| Area | Finding | Action |
|---|---|---|
| IndexArtworkJob | Already eager-loads relations used by toSearchableArray() |
NO CHANGE |
| IndexUserJob | Review: not a heat source; search queue empty | NO CHANGE |
| Heat command | Was one giant hydrate, not N+1 | chunked load |
| Studio / Journey | Window helper issues bounded per-user SQL | NO CHANGE |
| Public artwork list | idx_artworks_browse; no new eager-load without evidence |
NO CHANGE |
| Academy / forum cards | No production digest proving N+1 | NO CHANGE (do not globally eager-load) |
10. Artwork list / detail
Public browse uses idx_artworks_browse (EXPLAIN backward index scan). Category/tag pivots have supporting indexes from batch1. NO CHANGE to list indexes.
11. Rankings / leaderboards
RankBuildScopeListsJob reads rank_artwork_scores with LIMIT 200 after M7.1 uniqueness. DB side is cheap relative to heat/snapshots. leaderboards:refresh 30d GROUP BY remains the expensive reader as the snapshot table grows — keep current SQL; do not FORCE INDEX yet. Do not redesign queue architecture.
12. Metric-snapshot findings
Table ~8.87M rows, ~1855 MB (data ~740 / index ~1115). Unique (artwork_id, bucket_hour).
Indexes:
| Index | Columns | Verdict |
|---|---|---|
| PRIMARY | id |
keep |
uq_artwork_bucket UNIQUE |
artwork_id, bucket_hour |
keep — upsert + per-artwork window |
idx_artwork_bucket |
artwork_id, bucket_hour |
exact duplicate of unique — DROP candidate |
idx_bucket_hour |
bucket_hour |
left prefix of idx_bucket_artwork — do not drop in this milestone |
idx_bucket_artwork |
bucket_hour, artwork_id |
keep — heat DISTINCT covering |
Do not reduce 30-day retention. Do not OPTIMIZE TABLE. Do not partition.
13. Forum / collections
forum:scan-posts --limit=250stays synchronous. No new index (PRIMARY + limit). Do not re-enable forum queues.- Collections maintenance queue remains disabled.
collections:sync-lifecycleis cheap date scans. Every-10-minute sync is fine from DB.
14. Academy
No P_S digest pointing at Academy. No speculative indexes or global eager-load.
15. Search / recommendations
IndexArtworkJobeager-load is sufficient (no internal N+1 of the listed relations).- Tag IDF GROUP BY ran once per batch job; cached 1h locally.
- Do not create a rec snapshot table (M4).
16. Table / index inventory (production estimates)
| Table | Rows (approx) | Notes |
|---|---|---|
artwork_metric_snapshots_hourly |
8.87M | index > data; duplicate secondary |
artworks |
(browse index used) | idx_artworks_browse |
artwork_stats |
1:1 artwork | PK artwork_id; heat upsert |
artwork_tag |
covering tag_id | IDF GROUP BY |
| forum posts | ~13k | scan LIMIT 250 |
| rec tables | unique artwork+type+version | M7.1 intact |
17. Proposed indexes
None for add. Equality+range combinations already have uq_artwork_bucket and idx_bucket_artwork.
18. Proposed redundant index removals
DROP idx_artwork_bucket only.
Migration: database/migrations/2026_08_23_120000_drop_duplicate_artwork_metric_snapshot_index.php
Do not run on production in the application deploy. Separate schema window.
Percona 8.4: secondary index drop is typically ALGORITHM=INPLACE, LOCK=NONE. Disk: reclaim is InnoDB-internal, not filesystem (same as M5 prune). Rollback: re-add the duplicate index (unnecessary).
idx_bucket_hour: NO CHANGE until a dedicated EXPLAIN proves no leftover readers after drop of the duplicate.
19. Write-amplification
~50k upserts/hour. Each extra secondary index updates on every write.
- Duplicate
idx_artwork_bucket: 100% write cost, 0% extra read benefit. - Remaining secondaries:
idx_bucket_artworkis justified by heat DISTINCT covering.
20. Pagination / COUNT
No production evidence of OFFSET thousands on artworks browse. Do not rewrite pagination. Exact Studio counts remain M5.1 semantics.
21. Lock / transaction findings
Row lock waits exist but average 4 ms. Heat upsert is keyed by artwork_id (not a single hot row). No isolation-level change. Heat no longer loads 2.4M rows into PHP, which reduces transaction/memory pressure around upserts.
22. MySQL config verdict
NO CHANGE to innodb_buffer_pool_size, max_connections, redo, tmp_table_size, table caches, optimizer. Pool not under pressure.
23. Ranked findings
| Severity | Item |
|---|---|
| HIGH | Heat command hydrating full 24h snapshot set — fixed (chunked SELECT + narrower columns) |
| MEDIUM | Tag IDF GROUP BY per rec batch — fixed (Cache::remember 1h) |
| MEDIUM | Duplicate idx_artwork_bucket write amp — migration created, not executed on prod |
| MEDIUM | 30d leaderboard GROUP BY ~9M rows growing to ~36M — monitor; no FORCE INDEX yet |
| LOW | Sort_merge_passes / Select_scan cumulative — relate to heat DISTINCT temp; re-check after heat deploy |
| NO CHANGE | FPM max_children=14, Horizon counts, Redis, 30d retention, partition, OPTIMIZE, MySQL knobs, forum queues, rec snapshot table |
24. Files changed
app/Console/Commands/RecalculateHeatCommand.phpapp/Jobs/RecComputeSimilarByTagsJob.phptests/Feature/Metrics/RecalculateHeatCommandTest.phptests/Feature/Recommendations/RecComputeSimilarJobsTest.phpdatabase/migrations/2026_08_23_120000_drop_duplicate_artwork_metric_snapshot_index.phpdocs/optimization-m10-database-query-optimization.md
25. Migrations created
2026_08_23_120000_drop_duplicate_artwork_metric_snapshot_index.php — local only until a separate production schema deploy.
26. Tests / results
pest tests/Feature/Metrics/RecalculateHeatCommandTest.php PASS
phpunit --filter tag_idf RecComputeSimilarJobsTest OK
27. Expected benefit
- Heat: peak PHP memory and rows hydrated drop from ~all 24h snapshots (~1–2M rows) to
--chunk× hours (default 1000 × 25). - Tag jobs: one
artwork_tagGROUP BY per hour per model version instead of once per batch/chain hop. - Duplicate index drop (later): less write/storage on 50k rows/hour; index size currently already > data.
28. Application deployment sequence
- Deploy code (heat chunking + IDF cache). Do not run the DROP INDEX migration.
- After one
nova:recalculate-heatcycle: confirm command duration and RSS vs previous (scheduler runtime). - After rec tags chain: confirm
artwork_tagGROUP BY frequency (slow log / P_S if grants added later). - FPM/Horizon/Redis unchanged.
29. Schema deployment sequence (separate)
- Confirm
SHOW INDEX FROM artwork_metric_snapshots_hourlystill showsuq_artwork_bucketandidx_artwork_bucketidentical columns. ALTER TABLE ... DROP INDEX idx_artwork_bucketvia the migration in a dedicated window (INPLACE, no app coupling).- Re-EXPLAIN heat DISTINCT — expect
idx_bucket_artworkunchanged. - Do not drop
idx_bucket_hourin the same change.
30. Rollback
- App: revert heat/IDF commits; heat semantics unchanged (same formula, same upsert keys).
- Cache:
Cache::forget('rec:tag-idf:'.$modelVersion)or wait TTL. - Index: re-create
idx_artwork_bucketonly if a reader somehow used the name (none found; unique covers the same left prefix).
Production verification plan (per fix)
| Fix | Verify |
|---|---|
| Heat chunking | scheduler runtime; rows examined still on covering index; PHP memory; artwork_stats.heat_score still updates |
| IDF cache | query count of artwork_tag GROUP BY across a tags job chain |
| DROP duplicate index (later) | EXPLAIN heat + Studio window + prune WHERE bucket_hour < before/after |