Ship production optimization M1-M12.5A: queues, metrics, HTTP observability, and vector search reliability.
Keep similar-ai from tripping the global circuit on a lone URL 502, clamp Qdrant search to 100, and add Server-Timing plus slow-request logging. Studio shared props, Academy S3 exists caching, heat chunking, and Redis/scheduler hygiene stay in this rollout.
This commit is contained in:
@@ -39,6 +39,8 @@ The example worker listens on the application queues currently used by Skinbase
|
||||
search,forum-security,forum-moderation,vision,recommendations,discovery,mail,default
|
||||
```
|
||||
|
||||
Production uses **Horizon**, not this standalone `queue:work` example. Horizon supervisors must list every queue that application code pushes to. The `mail` queue is consumed by `supervisor-mail` in `config/horizon.php` (isolated from search/recommendation workers).
|
||||
|
||||
If you split workloads across dedicated workers, make sure the Scout/Meilisearch `search` queue and any queues configured through `VISION_QUEUE`, `RECOMMENDATIONS_QUEUE`, or `DISCOVERY_QUEUE` are explicitly covered by at least one worker process.
|
||||
|
||||
To use it on a Debian/Ubuntu server:
|
||||
|
||||
@@ -178,7 +178,7 @@ Examples below are representative. For the full option list of any Artisan comma
|
||||
| Command | Why it is used | Example |
|
||||
| --- | --- | --- |
|
||||
| `nova:metrics-snapshot-hourly` | Collect hourly metric snapshots used by rising/heat calculations | `php artisan nova:metrics-snapshot-hourly --days=1` |
|
||||
| `nova:prune-metric-snapshots` | Delete old hourly metric snapshots past the retention window | `php artisan nova:prune-metric-snapshots --keep-days=90` |
|
||||
| `nova:prune-metric-snapshots` | Delete old hourly metric snapshots past the retention window (batched) | `php artisan nova:prune-metric-snapshots --keep-days=30 --chunk=5000` |
|
||||
| `nova:recalculate-heat` | Recalculate heat and momentum scores for the rising engine | `php artisan nova:recalculate-heat --days=30 --chunk=500` |
|
||||
| `nova:recalculate-rankings` | Recalculate V2 ranking scores | `php artisan nova:recalculate-rankings --chunk=500 --sync-rank-scores` |
|
||||
|
||||
|
||||
@@ -0,0 +1,290 @@
|
||||
# M1 — Nightly recommendation job failures
|
||||
|
||||
## 1. Executive Summary
|
||||
|
||||
Production Horizon marks `RecComputeSimilarByBehaviorJob` and `RecComputeSimilarHybridJob` failed every night with `MaxAttemptsExceededException` at **~02:18** and **~02:31**.
|
||||
|
||||
This is **not** the old 90-second Supervisor `queue:work` timeout. Production uses Horizon workers with **timeout=960**. The Redis queue connection has **`retry_after=90`**. A still-running nightly rebuild is treated as abandoned, re-queued, and Horizon `--tries=1` fails it.
|
||||
|
||||
```text
|
||||
ROOT-H — Redis retry_after (90s) shorter than job runtime, interacting with Horizon tries=1
|
||||
and a single-job full-catalog scan (~50k artworks).
|
||||
```
|
||||
|
||||
Local fix (not deployed):
|
||||
|
||||
1. Default `queue.connections.redis.retry_after` **1080** (> Horizon 960).
|
||||
2. Fan-out Behavior and Hybrid rebuilds into cursor batches of 200, matching `RecComputeSimilarByTagsJob` (already chunked in production).
|
||||
|
||||
Mail queue `LLEN=3` is **deferred** (not entangled).
|
||||
|
||||
---
|
||||
|
||||
## 2. Production Symptom
|
||||
|
||||
```text
|
||||
failed_jobs ≈ 329
|
||||
RecComputeSimilarByBehaviorJob ~02:18 daily
|
||||
RecComputeSimilarHybridJob ~02:31 daily
|
||||
exception: Illuminate\Queue\MaxAttemptsExceededException
|
||||
"... has been attempted too many times."
|
||||
```
|
||||
|
||||
Schedule (Europe/Ljubljana):
|
||||
|
||||
| Job | dailyAt |
|
||||
| --- | ------- |
|
||||
| RecComputeSimilarByTagsJob | 02:00 |
|
||||
| RecComputeSimilarByBehaviorJob | 02:15 |
|
||||
| RecComputeSimilarHybridJob | 02:30 |
|
||||
|
||||
Tags does **not** appear in the nightly failure list (already cursor-batched).
|
||||
|
||||
---
|
||||
|
||||
## 3. Production Evidence
|
||||
|
||||
PRODUCTION SERVER (`ssh server3`, read-only):
|
||||
|
||||
- `config('queue.connections.redis.retry_after')` = **90**
|
||||
- `config('horizon.defaults.supervisor-default.timeout')` = **960**
|
||||
- Horizon worker: `--queue=search --timeout=960 --tries=1 --memory=128`
|
||||
- `recommendations.queue` = `default`
|
||||
- CLI `memory_limit` = `-1` (Horizon 128 MB is post-job worker recycle, not the kill mechanism at 90s)
|
||||
- Failed exception text is MaxAttemptsExceeded **without** a wrapped `TimeoutExceededException` / memory error
|
||||
- Timing:
|
||||
|
||||
```text
|
||||
Behavior 02:15 dispatch → fail ~02:18 (~180s or ~90s after the minute the scheduler actually queued it)
|
||||
Hybrid 02:30 dispatch → fail ~02:31:33 (≈ retry_after 90s)
|
||||
```
|
||||
|
||||
Laravel Redis queues use `retry_after` as the reservation TTL. If the job is still reserved after 90s, Redis exposes it again. The next pop has `attempts=2`. Horizon `--tries=1` then raises `MaxAttemptsExceededException` **before** the job body runs.
|
||||
|
||||
That matches the Hybrid `tries=1` failure ~90 seconds after start. Behavior’s class `$tries=2` is overridden by the worker `--tries=1`.
|
||||
|
||||
No production files, jobs, or services were changed during M1.
|
||||
|
||||
---
|
||||
|
||||
## 4. Local vs Production Code
|
||||
|
||||
| Item | Production `5af95f65` | Local `f52879ed` (before M1) |
|
||||
| ---- | --------------------- | ---------------------------- |
|
||||
| RecComputeSimilarByBehaviorJob | `$tries=2`, `$timeout=600`, **full `chunkById` catalog** | same |
|
||||
| RecComputeSimilarHybridJob | `$tries=1`, `$timeout=900`, **full catalog**, WithoutOverlapping only per-artwork | same |
|
||||
| RecComputeSimilarByTagsJob | cursor `afterArtworkId` batches of 200 | same |
|
||||
| redis retry_after | 90 | 90 |
|
||||
|
||||
Local did **not** already contain this fix. M1 implements it.
|
||||
|
||||
---
|
||||
|
||||
## 5. Job Architecture (after M1)
|
||||
|
||||
Nightly catalog rebuild (`artworkId=null`):
|
||||
|
||||
```text
|
||||
load ≤ batchSize artworks after cursor
|
||||
→ process each
|
||||
→ if batch full, dispatch next job with afterArtworkId = last id
|
||||
```
|
||||
|
||||
Per-artwork rebuilds (observers) unchanged: Hybrid still uses `WithoutOverlapping` + `dontRelease()`.
|
||||
|
||||
| Job | tries | timeout | batch |
|
||||
| --- | ----: | ------: | ----: |
|
||||
| Behavior | 2 | 120s | 200 |
|
||||
| Hybrid | 1 | 180s | 200 |
|
||||
| Tags | 2 | 600s | 200 (already) |
|
||||
|
||||
---
|
||||
|
||||
## 6. Scheduler Flow
|
||||
|
||||
Unchanged. `Schedule::job(...)->dailyAt(...)->withoutOverlapping()`.
|
||||
|
||||
Behavior and Hybrid still start at 02:15 / 02:30; work continues as chained queue jobs instead of one 15+ minute reservation.
|
||||
|
||||
Root crontab that invokes `schedule:run` remains unseen (sudo). Empirical evidence (nightly fails + 10:30 sitemap files) shows the scheduler **is** running.
|
||||
|
||||
---
|
||||
|
||||
## 7. Horizon Configuration
|
||||
|
||||
Production supervisors unchanged in M1 (no deploy).
|
||||
|
||||
Local comment on `supervisor-default.timeout` now documents that **`retry_after` must exceed 960**.
|
||||
|
||||
Horizon `tries=1` remains. Batch jobs should finish well under 90s; `retry_after=1080` is the safety net for any remaining long job.
|
||||
|
||||
---
|
||||
|
||||
## 8. Failed Job Analysis
|
||||
|
||||
Grouped: Behavior then Hybrid, every night, same exception. No payload dump. No evidence of immediate SQL exception (would usually wrap a QueryException). No SIGKILL/memory lines inspected in failed_jobs. Horizon log at inspection time showed healthy `IndexUserJob` / Scout jobs, not the 02:xx window.
|
||||
|
||||
---
|
||||
|
||||
## 9. Lock / Middleware Analysis
|
||||
|
||||
Full-catalog Behavior: **no** WithoutOverlapping.
|
||||
|
||||
Full-catalog Hybrid: `middleware()` returns `[]` when `artworkId === null`. **Not** the failure path.
|
||||
|
||||
Scheduler `withoutOverlapping()` would **skip** dispatch, not fail a job.
|
||||
|
||||
**Disproved:** `$tries=1` + WithoutOverlapping release as the nightly cause.
|
||||
|
||||
**Proved:** Redis reservation TTL 90s vs long-running catalog job.
|
||||
|
||||
---
|
||||
|
||||
## 10. Runtime / Memory Analysis
|
||||
|
||||
PHP CLI memory is unlimited. Horizon `--memory=128` typically exits **after** a job. A 50k-artwork in-memory scan could still grow, but the 90-second fail clock is Redis, not 128 MB.
|
||||
|
||||
After chunking, each batch holds 200 models + pair rows, similar to the tags job that already succeeds.
|
||||
|
||||
---
|
||||
|
||||
## 11. Query Analysis
|
||||
|
||||
Behavior, per artwork:
|
||||
|
||||
- `UNION` on `rec_item_pairs` (`a_artwork_id` / `b_artwork_id` indexed)
|
||||
- `artworks` `whereIn` related ids
|
||||
- `updateOrCreate` on `rec_artwork_recs`
|
||||
|
||||
~50k × 3 queries in **one** reservation was unbounded vs 90s.
|
||||
|
||||
No new index in M1. Chunking reduces reservation time; query shape unchanged.
|
||||
|
||||
---
|
||||
|
||||
## 12. Root Cause Classification
|
||||
|
||||
```text
|
||||
ROOT-H — Redis retry_after=90s < running catalog job
|
||||
+ Horizon --tries=1
|
||||
+ Behavior/Hybrid process the full catalog in one job
|
||||
```
|
||||
|
||||
Not ROOT-C (memory) as primary. Not ROOT-B (900s timeout) as the **observed** fail time (~90s). After fixing retry_after alone, ROOT-B could appear next (~600–900s). Chunking prevents that.
|
||||
|
||||
---
|
||||
|
||||
## 13. Fix
|
||||
|
||||
1. `config/queue.php` redis `retry_after` default **1080** (`REDIS_QUEUE_RETRY_AFTER`).
|
||||
2. `.env.example` documents the production requirement.
|
||||
3. Behavior + Hybrid catalog runs use the same cursor fan-out as Tags.
|
||||
4. Structured batch logs: `duration_ms`, `memory_mb`, `processed`, `has_more` (no PII).
|
||||
5. Per-artwork try/catch on Behavior (Hybrid already had it).
|
||||
|
||||
Correctness: same scoring/persist per artwork; batches are sequential by `id`. Idempotent `updateOrCreate`.
|
||||
|
||||
---
|
||||
|
||||
## 14. Files Changed
|
||||
|
||||
```text
|
||||
config/queue.php
|
||||
config/horizon.php
|
||||
.env.example
|
||||
app/Jobs/RecComputeSimilarByBehaviorJob.php
|
||||
app/Jobs/RecComputeSimilarHybridJob.php
|
||||
tests/Unit/Jobs/RecComputeSimilarByBehaviorJobTest.php
|
||||
tests/Unit/Jobs/RecComputeSimilarHybridJobTest.php
|
||||
tests/Feature/Recommendations/RecComputeSimilarJobsTest.php
|
||||
docs/optimization-m1-recommendation-jobs.md
|
||||
```
|
||||
|
||||
No migrations.
|
||||
|
||||
---
|
||||
|
||||
## 15. Tests
|
||||
|
||||
```text
|
||||
php artisan test --filter=RecComputeSimilar
|
||||
Tests: 13 passed (16 assertions)
|
||||
Duration: 15.67s
|
||||
```
|
||||
|
||||
Covers retry_after vs Horizon timeout, cursor dispatch, partial batch no-dispatch, hybrid idempotency.
|
||||
|
||||
Full suite not re-run (known unrelated failures from earlier audits).
|
||||
|
||||
---
|
||||
|
||||
## 16. Performance Impact
|
||||
|
||||
Expected: many short default-queue jobs (~50k/200 ≈ 250 Behavior + 250 Hybrid) instead of two 15-minute jobs. Worker slots free between batches. Redis no longer duplicates in-flight catalog jobs.
|
||||
|
||||
---
|
||||
|
||||
## 17. Correctness Impact
|
||||
|
||||
Same rec lists; order of artwork processing remains increasing `id`. A crash mid-chain can resume by re-dispatching from the last completed cursor (or wait for the next night). No silent catch-all.
|
||||
|
||||
---
|
||||
|
||||
## 18. Deployment Requirements
|
||||
|
||||
```text
|
||||
1. Deploy this release (includes config/queue.php).
|
||||
2. Set production .env REDIS_QUEUE_RETRY_AFTER=1080 if you pin the value
|
||||
(default in config is already 1080 if unset).
|
||||
3. php artisan config:cache (normal deploy optimize).
|
||||
4. Restart Horizon via the usual Supervisor restart / queue:restart.
|
||||
```
|
||||
|
||||
Without restarting Horizon, workers keep old cached config.
|
||||
|
||||
---
|
||||
|
||||
## 19. Post-Deploy Verification
|
||||
|
||||
```text
|
||||
1. php artisan horizon:status → running
|
||||
2. php artisan tinker: config('queue.connections.redis.retry_after') === 1080
|
||||
3. Next night: no new MaxAttemptsExceeded for RecComputeSimilarByBehavior/Hybrid
|
||||
4. rec_artwork_recs computed_at advances for similar_behavior and similar_hybrid
|
||||
5. Horizon log: "[RecComputeSimilarByBehavior] Batch complete" with has_more true then false
|
||||
6. failed_jobs count not increasing for these classes
|
||||
7. Optional: sample duration_ms / memory_mb
|
||||
```
|
||||
|
||||
Do not retry the 329 historical failures unless product wants a one-off backfill; they are the old reservation pattern.
|
||||
|
||||
---
|
||||
|
||||
## 20. Rollback Plan
|
||||
|
||||
```text
|
||||
- Roll back the release symlink / previous release
|
||||
- config:cache + Horizon restart
|
||||
- No DB migration to roll back
|
||||
- Cursor jobs in flight may still complete; safe
|
||||
- No Redis lock cleanup required for catalog jobs
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 21. Deferred Findings
|
||||
|
||||
```text
|
||||
DEFER TO M2/M1B: queues:mail LLEN=3, Horizon does not consume `mail`
|
||||
DEFER: sitemap index stale
|
||||
DEFER: artwork_metric_snapshots_hourly size
|
||||
DEFER: swap / FPM / Sentry sample rates
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 22. Remaining Unknowns
|
||||
|
||||
- Exact `schedule:run` crontab user (sudo required)
|
||||
- Whether any Behavior run ever reached the 600s job timeout (masked by 90s retry_after)
|
||||
- Historical failed_jobs payloads not inspected (by design)
|
||||
@@ -0,0 +1,290 @@
|
||||
# M10 — Database Query & Slow Query Optimization
|
||||
|
||||
```text
|
||||
STATUS: COMPLETE (read-only production audit + two local application fixes)
|
||||
PRODUCTION: no ALTER / DROP / OPTIMIZE / MySQL config / restart
|
||||
WHEN: 2026-08-23
|
||||
```
|
||||
|
||||
Production investigation used `ssh server3` as Skinbase (`skinbase@localhost`). Slow log file and `performance_schema.events_statements_summary_by_digest` were **not readable** with this grant. Cost ranking is from `SHOW GLOBAL STATUS`, `SHOW TABLE STATUS`, `SHOW INDEX`, `EXPLAIN`, and Laravel source mapping.
|
||||
|
||||
---
|
||||
|
||||
## 1. Production DB health snapshot
|
||||
|
||||
| Metric | Value |
|
||||
| ------ | ----- |
|
||||
| Version | Percona Server 8.4.11-11 |
|
||||
| Uptime | ~122233 s (~34 h) |
|
||||
| Threads_connected / Threads_running | 12 / 2 |
|
||||
| max_connections / Max_used_connections | 120 / 26 |
|
||||
| Slow_queries (cumulative) | 14301 (`long_query_time=0.5`, `min_examined_row_limit=100`) |
|
||||
| Created_tmp_disk_tables | 21 (low vs memory temps) |
|
||||
| Select_full_join | 30772 |
|
||||
| Select_scan | ~1.15M |
|
||||
| Sort_merge_passes | 107022 |
|
||||
| InnoDB buffer pool | 10 GB; hit ratio ~99.9996%; pages_data ~24% |
|
||||
| Innodb_row_lock_waits | 40794; avg ~4 ms; max ~449 ms |
|
||||
|
||||
Buffer pool and connection headroom are healthy. Cumulative `Slow_queries` is **not** treated as proof of a current problem (uptime + historical counter).
|
||||
|
||||
---
|
||||
|
||||
## 2. Slow-log configuration
|
||||
|
||||
| Variable | Production |
|
||||
| -------- | ---------- |
|
||||
| slow_query_log | ON |
|
||||
| slow_query_log_file | `/var/log/mysql/slow.log` |
|
||||
| long_query_time | 0.5 |
|
||||
| min_examined_row_limit | 100 |
|
||||
| log_queries_not_using_indexes | (not changed; not used as a tuning lever) |
|
||||
|
||||
**Do not change** these during M10. The log file is **not readable** as the app user. `pt-query-digest` was not used.
|
||||
|
||||
---
|
||||
|
||||
## 3–4. Slow-log / Performance Schema top fingerprints
|
||||
|
||||
**Unavailable** (`SELECT` denied on `events_statements_summary_by_digest`; slow.log unreadable).
|
||||
|
||||
Substitutes: `EXPLAIN` on known high-volume SQL + code-path frequency.
|
||||
|
||||
---
|
||||
|
||||
## 5–7. Top queries by total latency / average / rows examined
|
||||
|
||||
Ranked by **evidence from EXPLAIN + schedule frequency**, not P_S.
|
||||
|
||||
| Fingerprint | Feature | Path | rows examined (EXPLAIN) | Index | Severity |
|
||||
| ----------- | ------- | ---- | ----------------------- | ----- | -------- |
|
||||
| DISTINCT `artwork_id` WHERE `bucket_hour` BETWEEN ? AND ? | Rising heat | scheduled `nova:recalculate-heat` | ~2.46M covering `idx_bucket_artwork` | Using index + temp | HIGH (was hydrating all rows in PHP) |
|
||||
| SELECT snapshots WHERE `bucket_hour` BETWEEN AND `artwork_id` IN (all IDs) | Rising heat | same command (pre-fix) | ~2.46M | `idx_bucket_artwork` | HIGH — **fixed locally** (chunked) |
|
||||
| GROUP BY `artwork_id` 30d window | monthly leaderboards | `LeaderboardService` / `ArtworkHourlySnapshotWindow` | ~9.12M type=index `uq_artwork_bucket` | unique | MEDIUM (grows with 30d table) |
|
||||
| `artwork_tag` GROUP BY `tag_id` COUNT | rec tags IDF | `RecComputeSimilarByTagsJob::handle` every batch | ~112k covering | `artwork_tag_tag_id_index` | MEDIUM — **fixed locally** (cache) |
|
||||
| public list `is_public`+`is_approved` ORDER BY | HTTP browse | artwork listing | uses `idx_artworks_browse` backward | OK | NO CHANGE |
|
||||
| eligible snapshot IDs | hourly snapshot | `MetricsSnapshotHourlyCommand` | `artworks_is_approved_index` + stats PK | OK | NO CHANGE |
|
||||
| forum `reviewed=false` LIMIT 250 | `forum:scan-posts` | PRIMARY range | 13k-scale table | OK | NO CHANGE |
|
||||
| collections lifecycle WHERE dates | `collections:sync-lifecycle` every 10m | small tables | OK | LOW / NO CHANGE |
|
||||
|
||||
---
|
||||
|
||||
## 8. Mapping queries → Laravel
|
||||
|
||||
| Query | Source |
|
||||
| ----- | ------ |
|
||||
| Heat DISTINCT + snapshot load + `artwork_stats` upsert | `app/Console/Commands/RecalculateHeatCommand.php::handle` |
|
||||
| Hourly eligible IDs + upsert | `app/Console/Commands/MetricsSnapshotHourlyCommand` |
|
||||
| 30d window MAX/MIN deltas | `app/Services/Metrics/ArtworkHourlySnapshotWindow.php` used by `StudioMetricsService`, `CreatorStudioOverviewService`, `LeaderboardService`, `CreatorJourneyService` |
|
||||
| Rank lists LIMIT 200 | `app/Jobs/RankBuildScopeListsJob::fetchCandidates` (`rank_artwork_scores`) |
|
||||
| Tag IDF GROUP BY | `app/Jobs/RecComputeSimilarByTagsJob::tagFrequencies` |
|
||||
| Search document | `app/Jobs/IndexArtworkJob::handle` (`with` user, group, tags, categories.contentType, stats, awardStat) |
|
||||
| Forum scan | `packages/klevze/Plugins/Forum` + schedule `forum:scan-posts --limit=250` |
|
||||
| Collections sync | `app/Console/Commands/SyncCollectionLifecycleCommand.php` |
|
||||
|
||||
---
|
||||
|
||||
## 9. N+1 findings
|
||||
|
||||
| Area | Finding | Action |
|
||||
| ---- | ------- | ------ |
|
||||
| IndexArtworkJob | Already eager-loads relations used by `toSearchableArray()` | NO CHANGE |
|
||||
| IndexUserJob | Review: not a heat source; search queue empty | NO CHANGE |
|
||||
| Heat command | Was one giant hydrate, not N+1 | chunked load |
|
||||
| Studio / Journey | Window helper issues bounded per-user SQL | NO CHANGE |
|
||||
| Public artwork list | `idx_artworks_browse`; no new eager-load without evidence | NO CHANGE |
|
||||
| Academy / forum cards | No production digest proving N+1 | NO CHANGE (do not globally eager-load) |
|
||||
|
||||
---
|
||||
|
||||
## 10. Artwork list / detail
|
||||
|
||||
Public browse uses `idx_artworks_browse` (EXPLAIN backward index scan). Category/tag pivots have supporting indexes from batch1. **NO CHANGE** to list indexes.
|
||||
|
||||
---
|
||||
|
||||
## 11. Rankings / leaderboards
|
||||
|
||||
`RankBuildScopeListsJob` reads `rank_artwork_scores` with LIMIT 200 after M7.1 uniqueness. DB side is cheap relative to heat/snapshots. `leaderboards:refresh` 30d GROUP BY remains the expensive reader as the snapshot table grows — keep current SQL; do **not** FORCE INDEX yet. Do **not** redesign queue architecture.
|
||||
|
||||
---
|
||||
|
||||
## 12. Metric-snapshot findings
|
||||
|
||||
Table ~8.87M rows, ~1855 MB (data ~740 / index ~1115). Unique `(artwork_id, bucket_hour)`.
|
||||
|
||||
Indexes:
|
||||
|
||||
| Index | Columns | Verdict |
|
||||
| ----- | ------- | ------- |
|
||||
| PRIMARY | `id` | keep |
|
||||
| `uq_artwork_bucket` UNIQUE | `artwork_id`, `bucket_hour` | keep — upsert + per-artwork window |
|
||||
| `idx_artwork_bucket` | `artwork_id`, `bucket_hour` | **exact duplicate of unique** — DROP candidate |
|
||||
| `idx_bucket_hour` | `bucket_hour` | left prefix of `idx_bucket_artwork` — **do not drop in this milestone** |
|
||||
| `idx_bucket_artwork` | `bucket_hour`, `artwork_id` | keep — heat DISTINCT covering |
|
||||
|
||||
Do **not** reduce 30-day retention. Do **not** `OPTIMIZE TABLE`. Do **not** partition.
|
||||
|
||||
---
|
||||
|
||||
## 13. Forum / collections
|
||||
|
||||
- `forum:scan-posts --limit=250` stays **synchronous**. No new index (PRIMARY + limit). Do not re-enable forum queues.
|
||||
- Collections maintenance **queue remains disabled**. `collections:sync-lifecycle` is cheap date scans. Every-10-minute sync is fine from DB.
|
||||
|
||||
---
|
||||
|
||||
## 14. Academy
|
||||
|
||||
No P_S digest pointing at Academy. No speculative indexes or global eager-load.
|
||||
|
||||
---
|
||||
|
||||
## 15. Search / recommendations
|
||||
|
||||
- `IndexArtworkJob` eager-load is sufficient (no internal N+1 of the listed relations).
|
||||
- Tag IDF GROUP BY ran **once per batch job**; cached 1h locally.
|
||||
- Do **not** create a rec snapshot table (M4).
|
||||
|
||||
---
|
||||
|
||||
## 16. Table / index inventory (production estimates)
|
||||
|
||||
| Table | Rows (approx) | Notes |
|
||||
| ----- | ------------- | ----- |
|
||||
| `artwork_metric_snapshots_hourly` | 8.87M | index > data; duplicate secondary |
|
||||
| `artworks` | (browse index used) | `idx_artworks_browse` |
|
||||
| `artwork_stats` | 1:1 artwork | PK `artwork_id`; heat upsert |
|
||||
| `artwork_tag` | covering tag_id | IDF GROUP BY |
|
||||
| forum posts | ~13k | scan LIMIT 250 |
|
||||
| rec tables | unique artwork+type+version | M7.1 intact |
|
||||
|
||||
---
|
||||
|
||||
## 17. Proposed indexes
|
||||
|
||||
**None** for add. Equality+range combinations already have `uq_artwork_bucket` and `idx_bucket_artwork`.
|
||||
|
||||
---
|
||||
|
||||
## 18. Proposed redundant index removals
|
||||
|
||||
**DROP `idx_artwork_bucket`** only.
|
||||
|
||||
Migration: `database/migrations/2026_08_23_120000_drop_duplicate_artwork_metric_snapshot_index.php`
|
||||
|
||||
**Do not run on production in the application deploy.** Separate schema window.
|
||||
|
||||
Percona 8.4: secondary index drop is typically `ALGORITHM=INPLACE`, `LOCK=NONE`. Disk: reclaim is InnoDB-internal, not filesystem (same as M5 prune). Rollback: re-add the duplicate index (unnecessary).
|
||||
|
||||
`idx_bucket_hour`: **NO CHANGE** until a dedicated EXPLAIN proves no leftover readers after drop of the duplicate.
|
||||
|
||||
---
|
||||
|
||||
## 19. Write-amplification
|
||||
|
||||
~50k upserts/hour. Each extra secondary index updates on every write.
|
||||
|
||||
- Duplicate `idx_artwork_bucket`: **100% write cost, 0% extra read benefit**.
|
||||
- Remaining secondaries: `idx_bucket_artwork` is justified by heat DISTINCT covering.
|
||||
|
||||
---
|
||||
|
||||
## 20. Pagination / COUNT
|
||||
|
||||
No production evidence of OFFSET thousands on artworks browse. Do **not** rewrite pagination. Exact Studio counts remain M5.1 semantics.
|
||||
|
||||
---
|
||||
|
||||
## 21. Lock / transaction findings
|
||||
|
||||
Row lock waits exist but average 4 ms. Heat upsert is keyed by `artwork_id` (not a single hot row). **No isolation-level change.** Heat no longer loads 2.4M rows into PHP, which reduces transaction/memory pressure around upserts.
|
||||
|
||||
---
|
||||
|
||||
## 22. MySQL config verdict
|
||||
|
||||
**NO CHANGE** to `innodb_buffer_pool_size`, `max_connections`, redo, tmp_table_size, table caches, optimizer. Pool not under pressure.
|
||||
|
||||
---
|
||||
|
||||
## 23. Ranked findings
|
||||
|
||||
| Severity | Item |
|
||||
| -------- | ---- |
|
||||
| HIGH | Heat command hydrating full 24h snapshot set — **fixed** (chunked SELECT + narrower columns) |
|
||||
| MEDIUM | Tag IDF GROUP BY per rec batch — **fixed** (Cache::remember 1h) |
|
||||
| MEDIUM | Duplicate `idx_artwork_bucket` write amp — **migration created, not executed on prod** |
|
||||
| MEDIUM | 30d leaderboard GROUP BY ~9M rows growing to ~36M — monitor; no FORCE INDEX yet |
|
||||
| LOW | Sort_merge_passes / Select_scan cumulative — relate to heat DISTINCT temp; re-check after heat deploy |
|
||||
| NO CHANGE | FPM max_children=14, Horizon counts, Redis, 30d retention, partition, OPTIMIZE, MySQL knobs, forum queues, rec snapshot table |
|
||||
|
||||
---
|
||||
|
||||
## 24. Files changed
|
||||
|
||||
- `app/Console/Commands/RecalculateHeatCommand.php`
|
||||
- `app/Jobs/RecComputeSimilarByTagsJob.php`
|
||||
- `tests/Feature/Metrics/RecalculateHeatCommandTest.php`
|
||||
- `tests/Feature/Recommendations/RecComputeSimilarJobsTest.php`
|
||||
- `database/migrations/2026_08_23_120000_drop_duplicate_artwork_metric_snapshot_index.php`
|
||||
- `docs/optimization-m10-database-query-optimization.md`
|
||||
|
||||
---
|
||||
|
||||
## 25. Migrations created
|
||||
|
||||
`2026_08_23_120000_drop_duplicate_artwork_metric_snapshot_index.php` — **local only until a separate production schema deploy**.
|
||||
|
||||
---
|
||||
|
||||
## 26. Tests / results
|
||||
|
||||
```text
|
||||
pest tests/Feature/Metrics/RecalculateHeatCommandTest.php PASS
|
||||
phpunit --filter tag_idf RecComputeSimilarJobsTest OK
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 27. Expected benefit
|
||||
|
||||
- Heat: peak PHP memory and rows hydrated drop from ~all 24h snapshots (~1–2M rows) to `--chunk` × hours (default 1000 × 25).
|
||||
- Tag jobs: one `artwork_tag` GROUP BY per hour per model version instead of once per batch/chain hop.
|
||||
- Duplicate index drop (later): less write/storage on 50k rows/hour; index size currently already > data.
|
||||
|
||||
---
|
||||
|
||||
## 28. Application deployment sequence
|
||||
|
||||
1. Deploy code (heat chunking + IDF cache). **Do not** run the DROP INDEX migration.
|
||||
2. After one `nova:recalculate-heat` cycle: confirm command duration and RSS vs previous (scheduler runtime).
|
||||
3. After rec tags chain: confirm `artwork_tag` GROUP BY frequency (slow log / P_S if grants added later).
|
||||
4. FPM/Horizon/Redis unchanged.
|
||||
|
||||
---
|
||||
|
||||
## 29. Schema deployment sequence (separate)
|
||||
|
||||
1. Confirm `SHOW INDEX FROM artwork_metric_snapshots_hourly` still shows `uq_artwork_bucket` and `idx_artwork_bucket` identical columns.
|
||||
2. `ALTER TABLE ... DROP INDEX idx_artwork_bucket` via the migration in a dedicated window (INPLACE, no app coupling).
|
||||
3. Re-EXPLAIN heat DISTINCT — expect `idx_bucket_artwork` unchanged.
|
||||
4. Do not drop `idx_bucket_hour` in the same change.
|
||||
|
||||
---
|
||||
|
||||
## 30. Rollback
|
||||
|
||||
- App: revert heat/IDF commits; heat semantics unchanged (same formula, same upsert keys).
|
||||
- Cache: `Cache::forget('rec:tag-idf:'.$modelVersion)` or wait TTL.
|
||||
- Index: re-create `idx_artwork_bucket` only if a reader somehow used the name (none found; unique covers the same left prefix).
|
||||
|
||||
---
|
||||
|
||||
## Production verification plan (per fix)
|
||||
|
||||
| Fix | Verify |
|
||||
| --- | ------ |
|
||||
| Heat chunking | scheduler runtime; rows examined still on covering index; PHP memory; `artwork_stats.heat_score` still updates |
|
||||
| IDF cache | query count of `artwork_tag` GROUP BY across a tags job chain |
|
||||
| DROP duplicate index (later) | EXPLAIN heat + Studio window + prune `WHERE bucket_hour <` before/after |
|
||||
@@ -0,0 +1,322 @@
|
||||
# M11 — HTTP Endpoint Performance & Laravel Request Profiling
|
||||
|
||||
```text
|
||||
STATUS: COMPLETE (read-only production HTTP audit + local low-risk fixes)
|
||||
PRODUCTION: no cache flush, no restarts, no nginx/PHP/MySQL/FPM/Horizon changes
|
||||
WHEN: 2026-08-23 20:58–21:05 CEST
|
||||
```
|
||||
|
||||
Guest curls from `server3` to `https://skinbase.org` (HTTP/2, `Accept-Encoding: br`). Two sequential requests per path. No concurrency. No cache clears.
|
||||
|
||||
M10 snapshot indexes: **unchanged**. Duplicate-index DROP remains unrun; production does not have that index.
|
||||
|
||||
---
|
||||
|
||||
## 1. HTTP baseline table
|
||||
|
||||
Times are TTFB seconds. Size is **compressed** body (brotli) unless noted.
|
||||
|
||||
| Path | Status | 1st TTFB | 2nd TTFB | Size (br) | Cache-Control | Encoding |
|
||||
| ---- | ------ | -------- | -------- | --------- | ------------- | -------- |
|
||||
| `/` | 200 | 0.449 | 0.413 | ~22 KB | `public, max-age=60, s-maxage=300, swr=600` | br |
|
||||
| `/` uncompressed HTML | 200 | 0.444 | — | 286 KB | same | none (no AE) |
|
||||
| `/explore` | 200 | 0.504 | 0.490 | ~16 KB | `no-cache, private` | br |
|
||||
| `/discover/trending` | 200 | 0.449 | 0.434 | ~8.8 KB | `no-cache, private` | br |
|
||||
| `/uploads/latest` | 200 | 0.573 | 0.594 | ~11.5 KB | `no-cache, private` | br |
|
||||
| `/academy` | 200 | **1.283** | 0.838 | ~24 KB | `no-cache, private` | br |
|
||||
| `/academy/prompts` | 200 | **1.316** | **1.171** | ~29 KB | `no-cache, private` | br |
|
||||
| `/academy/prompts/popular` | 200 | 0.835 | — | — | private | br |
|
||||
| `/forum` | 200 | 0.512 | 0.525 | ~15 KB | `no-cache, private` | br |
|
||||
| `/collections` | **404** | 0.682 | 0.440 | ~8.7 KB | private | br |
|
||||
| `/collections/featured` | 200 | **1.037** | — | — | private | br |
|
||||
| `/rss/latest-uploads.xml` | 200 | 0.349 | 0.350 | ~13 KB | `no-cache, private` | br |
|
||||
| `/sitemap.xml` | 200 | 0.125 | 0.064 | 377 B | `public, max-age=21600` | br |
|
||||
| `/search?q=skin` | 200 | 0.467 | 0.520 | ~15 KB | private | br |
|
||||
| `/statistics` | **302 → /login** | 0.309 | 0.254 | 709 B | private | — |
|
||||
| `/leaderboard` | 200 | 0.439 | 0.435 | ~14 KB | private | br |
|
||||
| `/staff` | 200 | 0.382 | 0.413 | ~8 KB | private | br |
|
||||
| `/@killua` (profile) | 200 | **0.975** | **0.968** | 249 KB **uncompressed** | private | — |
|
||||
| `/favicon.ico` | 200 | 0.063 | — | — | static | — |
|
||||
|
||||
No `Server-Timing` header **before** this milestone. Added locally (`AddServerTiming`).
|
||||
|
||||
---
|
||||
|
||||
## 2. Warm vs second request
|
||||
|
||||
| Path | Δ | Interpretation |
|
||||
| ---- | - | -------------- |
|
||||
| `/` | −36 ms | Guest homepage `Cache::flexible` working |
|
||||
| `/sitemap.xml` | −61 ms | Public cache / file cache |
|
||||
| `/academy` | −445 ms | Application cache (Academy home payload) + SSR warm |
|
||||
| `/academy/prompts` | −145 ms | Still >1 s — listing + S3 `exists` + SSR |
|
||||
| Most others | ~0 | Cache already warm; remaining cost is PHP + SSR |
|
||||
|
||||
Did **not** flush Redis to manufacture cold tests.
|
||||
|
||||
---
|
||||
|
||||
## 3. Access-log latency analysis
|
||||
|
||||
`/etc/nginx/nginx.conf` uses **default** `access_log` (combined). **No custom `log_format`.**
|
||||
|
||||
Missing: `request_time`, `upstream_response_time`, `upstream_connect_time`.
|
||||
|
||||
Log file `/var/log/nginx/skinbase.org-access.log` is `www-data:adm` — app user **cannot read** it. **No p95/p99 from access logs.** Do not change nginx during M11.
|
||||
|
||||
---
|
||||
|
||||
## 4–5. Laravel logs / 5xx
|
||||
|
||||
`storage/logs/laravel.log` ~3.5 GB.
|
||||
|
||||
**CURRENT (hours around audit):** download warnings for missing originals (not 5xx). No Redis `LOADING`. No Predis memory exhaustion.
|
||||
|
||||
**RECENT (same day):** CLI errors (`nova:metrics-prune` not defined, tinker parse, Enhance worker unavailable). Not HTTP user-facing.
|
||||
|
||||
**HISTORICAL:** Redis LOADING / presence SMEMBERS / Predis OOM — M7/M8, not recurring in this window.
|
||||
|
||||
HTTP 5xx count from nginx: **unknown** (log unread + no timing format).
|
||||
|
||||
---
|
||||
|
||||
## 6. Middleware cost
|
||||
|
||||
Web append: `AddServerTiming` (new), `SecurityHeaders`, `RedirectLegacyProfileSubdomain`, `TrackOnlineVisitor`, `UpdateLastVisit`, `HandleInertiaRequests`, onboarding/email-upgrade.
|
||||
|
||||
| Middleware | Per request | Verdict |
|
||||
| ---------- | ----------- | ------- |
|
||||
| TrackOnlineVisitor | Redis GET + SETEX + SADD after response; bots included; RSS/sitemaps **now skipped** | LOW after skip |
|
||||
| UpdateLastVisit | DB update throttled 300s, auth only | NO CHANGE |
|
||||
| HandleInertiaRequests | auth user + **was** full `studioOptionsForUser` on every auth Inertia page | **HIGH for auth** — gated to `/studio` |
|
||||
| Session | skipped for some paths via ConditionalStartSession | NO CHANGE |
|
||||
|
||||
---
|
||||
|
||||
## 7. Inertia shared props
|
||||
|
||||
`app/Http/Middleware/HandleInertiaRequests.php::share`
|
||||
|
||||
| Prop | Was | Now |
|
||||
| ---- | --- | --- |
|
||||
| `auth.user` | profile lazy-load | `load('profile')` once |
|
||||
| `flash` | closures | same |
|
||||
| `cdn` / `features` | config | same |
|
||||
| `studio_groups` | **eager Group::with(owner.profile, members) + whereHas for every logged-in Inertia page** | **only when path starts with `studio`** |
|
||||
|
||||
`studio_groups` is consumed only by `resources/js/Layouts/StudioLayout.jsx`.
|
||||
|
||||
---
|
||||
|
||||
## 8. SSR audit
|
||||
|
||||
- Enabled: `config/inertia.php` `ssr.enabled=true`, `http://127.0.0.1:13714`
|
||||
- Process: `node …/bootstrap/ssr/ssr.js` (www-data)
|
||||
- Used for Inertia HTML first visits (Academy, forum, leaderboard, collections, profile)
|
||||
- Home `/` is **Blade**, not SSR
|
||||
- No SSR duration in logs before Server-Timing
|
||||
- Do **not** disable SSR globally
|
||||
- Academy/collections ~1 s likely **SSR + PHP**; second academy hit faster (cache)
|
||||
|
||||
---
|
||||
|
||||
## 9. Controller mapping (slowest)
|
||||
|
||||
| Path | Controller |
|
||||
| ---- | ---------- |
|
||||
| `/` | `App\Http\Controllers\Web\HomeController::index` → `HomepageService::all` / `allForUser` |
|
||||
| `/explore` | `ExploreController::index` (Meilisearch) |
|
||||
| `/discover/trending` | `DiscoverController::trending` |
|
||||
| `/academy` | `Academy\AcademyHomeController::index` |
|
||||
| `/academy/prompts` | `Academy\AcademyPromptController::index` |
|
||||
| `/leaderboard` | `LeaderboardPageController` → cached `LeaderboardService::getLeaderboard` |
|
||||
| `/collections/featured` | `CollectionDiscoveryController::featured` (6 discovery queries + SSR) |
|
||||
| `/search` | `SearchController::index` → Meilisearch + groups + news LIKE |
|
||||
| `/@user` | `User\ProfileController::showByUsername` |
|
||||
| `/download/artwork/{id}` | `ArtworkDownloadController` |
|
||||
| RSS | `RssFeedController::latestUploads` |
|
||||
|
||||
---
|
||||
|
||||
## 10–11. Query counts / M10 overlap
|
||||
|
||||
Local Pest: leaderboard as auth user **0** `groups` / `group_members` queries after the shared-prop gate.
|
||||
|
||||
Production EXPLAIN not repeated (M10). HTTP-facing:
|
||||
|
||||
- Leaderboard HTTP uses **Cache::remember**, not the 9M snapshot GROUP BY
|
||||
- Heat job is scheduler, not HTTP
|
||||
- Browse `/explore` uses Meilisearch then hydrate — TTFB ~500 ms
|
||||
|
||||
---
|
||||
|
||||
## 12–13. Redis / presence
|
||||
|
||||
Per tracked GET: `GET` record, `SETEX` 300s, `SADD` index. No `SMEMBERS` on the request path (M8 SSCAN prune remains).
|
||||
|
||||
RSS/sitemaps/robots **were tracked**; now skipped. Bots still tracked on HTML (by design).
|
||||
|
||||
---
|
||||
|
||||
## 14. Payload
|
||||
|
||||
Homepage uncompressed ~286 KB HTML (Inertia-like JSON in Blade). Brotli ~22 KB.
|
||||
|
||||
Profile `/@killua` 249 KB uncompressed HTML — large Inertia SSR document. MEDIUM, no rewrite this milestone.
|
||||
|
||||
Academy compressed 24–29 KB — fine; TTFB is CPU/S3/SSR not transfer.
|
||||
|
||||
---
|
||||
|
||||
## 15–21. Surface findings
|
||||
|
||||
**Home:** guest flexible cache 5s/30s fresh + 1800s stale. TTFB ~410–450 ms. Personalized `allForUser` uncached (not measured; private).
|
||||
|
||||
**Browse:** `/explore` ~500 ms; `no-cache, private` even for guests (CDN cannot cache). MEDIUM documentation only — personalization/maturity may require private.
|
||||
|
||||
**Artwork detail:** not sampled with a canonical slug this pass (explore HTML is CDN image URLs). Code path `ArtworkController::show` / web artwork.
|
||||
|
||||
**Search:** 467 ms including Meilisearch + optional news `LIKE` on content. Malformed query guard already in `SearchController`.
|
||||
|
||||
**Statistics:** auth-gated 302. Rankings `/leaderboard` 439 ms from cache.
|
||||
|
||||
**Profile:** ~970 ms both hits — SSR + heavy profile assembly. Documented, no rewrite.
|
||||
|
||||
**Academy:** slowest public family. Causes: Inertia SSR + `Storage::exists` on S3 for prompt preview variants + `courseLessons()->count()` when `lessons_count_cache` is 0 (`?:` bug). Fixes: Redis-remember exists 1h; use cache column without extra count.
|
||||
|
||||
**Forum:** 512 ms. Existing query-bound tests. Queues stay off.
|
||||
|
||||
**RSS:** 350 ms, two queries (`Artwork::published()->with(user)->limit`). Redis not required for generation. Presence writes removed from RSS.
|
||||
|
||||
**Download:** production currently logs missing originals **after** incrementing counters — **fixed** (404 first).
|
||||
|
||||
---
|
||||
|
||||
## 22. External / storage
|
||||
|
||||
Academy list `Storage::disk(s3)->exists` per variant — **was synchronous S3 on HTTP**. Now cached 1h.
|
||||
|
||||
Download `File::isFile` local originals.
|
||||
|
||||
Enhance worker errors are CLI/job, not public pages.
|
||||
|
||||
---
|
||||
|
||||
## 23. Cache / stampede
|
||||
|
||||
Homepage already `flexible` (stale-while-revalidate). Leaderboard `Cache::remember` — stampede possible but 439 ms and scheduled refresh exist. **No new caches** except Academy asset-exists.
|
||||
|
||||
---
|
||||
|
||||
## 24. Bots / deep pagination
|
||||
|
||||
Access logs unreadable. Historical malformed search URLs remain historical. **No UA blocking.**
|
||||
|
||||
---
|
||||
|
||||
## 25. HTTP cache / compression / static
|
||||
|
||||
- gzip + brotli **on** in nginx (confirmed `content-encoding: br` when AE sent)
|
||||
- Static `favicon.ico` 63 ms, not Laravel
|
||||
- `/build/*` access_log off in vhost
|
||||
- Guest `/` already public CDN-friendly; most Inertia HTML `private, no-cache` (auth/session/SSR)
|
||||
|
||||
---
|
||||
|
||||
## 26. TTFB classification
|
||||
|
||||
| Endpoint | Dominant cost |
|
||||
| -------- | ------------- |
|
||||
| `/` | PHP bootstrap + cached payload + HTML serialize (~400 ms floor) |
|
||||
| `/explore`, `/discover/*` | PHP + Meilisearch + Blade/Inertia |
|
||||
| `/academy*` | **SSR + PHP + (was) S3 exists** |
|
||||
| `/collections/featured` | **SSR + multiple discovery SQL** |
|
||||
| `/@user` | **SSR + profile assembly** |
|
||||
| `/leaderboard` | PHP + Redis cache hit |
|
||||
| `/sitemap.xml` | static/cached XML |
|
||||
| `/rss/*` | PHP + small SQL |
|
||||
|
||||
---
|
||||
|
||||
## 27. Ranked findings
|
||||
|
||||
| Severity | Item |
|
||||
| -------- | ---- |
|
||||
| HIGH | Auth Inertia `studio_groups` full group graph every page — **fixed** (studio path only) |
|
||||
| HIGH | Download counters/events before missing-file 404 — **fixed** |
|
||||
| HIGH | Academy S3 `exists` on listing/home cards — **fixed** (1h cache) |
|
||||
| MEDIUM | Academy / collections / profile ~1 s (SSR). Do not disable SSR |
|
||||
| MEDIUM | Guest explore/discover `Cache-Control: private` — CDN cannot cache |
|
||||
| MEDIUM | `/collections` 404; real routes `/collections/featured` etc. |
|
||||
| LOW | Presence on RSS — **fixed skip** |
|
||||
| LOW | `lessons_count_cache ?: count()` extra SQL — **fixed** |
|
||||
| NO CHANGE | FPM 14, Horizon, Redis size, snapshot indexes, forum queues, MySQL knobs |
|
||||
|
||||
---
|
||||
|
||||
## 28. Files changed
|
||||
|
||||
- `app/Http/Middleware/HandleInertiaRequests.php`
|
||||
- `app/Http/Middleware/TrackOnlineVisitor.php`
|
||||
- `app/Http/Middleware/AddServerTiming.php` (new)
|
||||
- `bootstrap/app.php`
|
||||
- `app/Http/Controllers/ArtworkDownloadController.php`
|
||||
- `app/Services/Academy/AcademyAccessService.php`
|
||||
- `tests/Feature/Http/StudioGroupsSharedPropTest.php`
|
||||
- `tests/Feature/Http/ArtworkDownloadMissingFileTest.php`
|
||||
- `tests/Feature/Http/RssPresenceSkipTest.php`
|
||||
- `docs/optimization-m11-http-endpoint-performance.md`
|
||||
|
||||
---
|
||||
|
||||
## 29. Tests
|
||||
|
||||
```text
|
||||
pest StudioGroupsSharedPropTest PASS
|
||||
pest ArtworkDownloadMissingFileTest PASS
|
||||
pest RssPresenceSkipTest PASS
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 30. Production deployment plan
|
||||
|
||||
1. Deploy application code only.
|
||||
2. Do **not** run M10 DROP INDEX.
|
||||
3. Do **not** restart Redis/MySQL/FPM/Horizon except normal PHP-FPM reload for PHP.
|
||||
4. SSR process already running; no disable.
|
||||
|
||||
---
|
||||
|
||||
## 31. Post-deploy verification
|
||||
|
||||
From `server3`, sequential:
|
||||
|
||||
```bash
|
||||
curl -sS -D - -o /dev/null --max-time 20 -H 'Accept-Encoding: br' https://skinbase.org/academy | grep -iE 'HTTP/|server-timing|cache-control'
|
||||
curl -sS -o /dev/null -w 'academy ttfb:%{time_starttransfer}\n' -H 'Accept-Encoding: br' https://skinbase.org/academy
|
||||
curl -sS -o /dev/null -w 'prompts ttfb:%{time_starttransfer}\n' -H 'Accept-Encoding: br' https://skinbase.org/academy/prompts
|
||||
curl -sS -o /dev/null -w 'home ttfb:%{time_starttransfer}\n' -H 'Accept-Encoding: br' https://skinbase.org/
|
||||
```
|
||||
|
||||
Expect `Server-Timing: app;desc="Laravel";dur=…`. Academy TTFB should move toward the warm ~0.8 s band more often (asset-exists cache). Guest home should stay ~400–450 ms.
|
||||
|
||||
---
|
||||
|
||||
## 32. Rollback
|
||||
|
||||
Revert the listed PHP files. Presence skip and Server-Timing are additive. Download 404-before-increment is the correct semantics (do not roll back unless a download of a missing file must still increment stats — it should not).
|
||||
|
||||
---
|
||||
|
||||
## Measurement note (acceptance)
|
||||
|
||||
Public guest pages already under ~500 ms except Academy (~1.3 s), collections featured (~1.0 s), profile (~0.97 s).
|
||||
|
||||
| Path | BEFORE (prod guest) | AFTER (expected) |
|
||||
| ---- | ------------------- | ---------------- |
|
||||
| `/` | 0.45 / 0.41 s | same |
|
||||
| `/academy` | 1.28 / 0.84 s | fewer S3 round-trips; SSR remains |
|
||||
| `/academy/prompts` | 1.32 / 1.17 s | fewer S3 round-trips |
|
||||
| Auth non-studio Inertia | unmeasured (group graph every request) | **0 extra group SQL** (tested) |
|
||||
| Missing download | counters + 404 | **404 only** (tested) |
|
||||
@@ -0,0 +1,303 @@
|
||||
# M12 — Production Performance Observability & Slow Request Logging
|
||||
|
||||
```text
|
||||
STATUS: COMPLETE (audit + local artifacts; production nginx/FPM/env NOT modified)
|
||||
WHEN: 2026-08-24 16:46 CEST
|
||||
PRODUCTION: read-only audit only
|
||||
```
|
||||
|
||||
This milestone is **observability**. No application behavior changes except:
|
||||
|
||||
- Server-Timing now wraps the full HTTP kernel (prepend) and can add `ssr;dur=` when Inertia SSR runs
|
||||
- Slow-request logging is **off until** `HTTP_SLOW_REQUEST_LOG_ENABLED=true`
|
||||
|
||||
---
|
||||
|
||||
## 1. Current nginx access logging
|
||||
|
||||
| Item | Production |
|
||||
| ---- | ---------- |
|
||||
| nginx | **1.26.3** |
|
||||
| `log_format` | **none** (default `combined`) |
|
||||
| http default | `access_log /var/log/nginx/access.log;` |
|
||||
| Skinbase vhost | `access_log /var/log/nginx/skinbase.org-access.log;` (no format → combined) |
|
||||
| `access_log off` | static `css/js/images`, `/photo/*.jpg` |
|
||||
| Query strings | **yes** in combined `$request` |
|
||||
| Timing vars | **absent** |
|
||||
| SHA256 `nginx.conf` | `d1720f312b054a2007ffa731158c8bf4c9bd64bba35ba3971551f28cbabd2bd5` |
|
||||
| SHA256 `skinbase.org.conf` | `4e7e4aa29cb406ccedf3e42613997683a92f1c8e2786bd087adab6012dd4b26f` |
|
||||
|
||||
Multiple vhosts have **separate** logs. `http {}` includes `/etc/nginx/conf.d/*.conf`.
|
||||
|
||||
**Cloudflare** sits in front (`server: cloudflare`, `cf-cache-status: DYNAMIC`, `00-cloudflare-realip.conf`). Origin `request_time` is nginx-on-server3, not browser RTT.
|
||||
|
||||
---
|
||||
|
||||
## 2. Logrotate
|
||||
|
||||
`/etc/logrotate.d/nginx`:
|
||||
|
||||
```
|
||||
/var/log/nginx/*.log {
|
||||
daily
|
||||
rotate 14
|
||||
compress
|
||||
delaycompress
|
||||
create 0640 www-data adm
|
||||
postrotate: invoke-rc.d nginx rotate
|
||||
}
|
||||
```
|
||||
|
||||
A new `/var/log/nginx/skinbase-performance.log` **is already covered**. **No extra logrotate rule.**
|
||||
|
||||
---
|
||||
|
||||
## 3. Current FPM slowlog
|
||||
|
||||
Pool `/etc/php/8.4/fpm/pool.d/skinbase.conf` (SHA256 `a0211876…`):
|
||||
|
||||
| Setting | Value |
|
||||
| ------- | ----- |
|
||||
| `pm.max_children` | 14 (**do not change**) |
|
||||
| `request_terminate_timeout` | 120s |
|
||||
| `request_slowlog_timeout` | **10s** (already on) |
|
||||
| `slowlog` | `/var/log/php8.4-fpm-skinbase-slow.log` (~23 MB) |
|
||||
| `catch_workers_output` | yes |
|
||||
| `pm.status_path` | `/fpm-status` localhost only |
|
||||
|
||||
**FPM slowlog recommendation: NO CHANGE.** 10s is appropriate for rare stacks. Do not lower to 2s in the same deploy as nginx JSON logging.
|
||||
|
||||
`/fpm-status` and `/fpm-ping` are `allow 127.0.0.1; deny all`. No public status. No `stub_status`.
|
||||
|
||||
---
|
||||
|
||||
## 4. Existing Server-Timing
|
||||
|
||||
Live header (server3 curl `/`):
|
||||
|
||||
```
|
||||
server-timing: app;desc="Laravel";dur=96.9
|
||||
cache-control: max-age=60, public, s-maxage=300, stale-while-revalidate=600
|
||||
```
|
||||
|
||||
M11 middleware was **web-appended**, so it missed earlier kernel/session work (TTFB ~0.45 s vs Laravel ~110 ms). M12 **prepends** it globally so `app;dur` is closer to full Laravel time. Streamed/download: skip if `headers_sent()`. No SQL/Redis/cookies in the header.
|
||||
|
||||
---
|
||||
|
||||
## 5. Proposed performance log format
|
||||
|
||||
File: `/var/log/nginx/skinbase-performance.log`
|
||||
Format name: `skinbase_perf` `escape=json`
|
||||
Prepared: `deploy/nginx/conf.d/skinbase-perf-log-format.conf`
|
||||
Vhost extra line: `deploy/nginx/skinbase-vhost-performance-log.snippet`
|
||||
|
||||
Existing combined log **stays**.
|
||||
|
||||
Static locations with `access_log off` stay off (not forced through PHP).
|
||||
|
||||
`$upstream_*` is `"-"` for static; analyzer treats as null. Fastcgi PHP has numeric upstream times. `$upstream_cache_status` is typically empty (no nginx proxy_cache for Laravel) — still logged as `cache`.
|
||||
|
||||
---
|
||||
|
||||
## 6. Privacy / security
|
||||
|
||||
**Logged:** time, host, method, `$uri` (path only), status, timings, bytes, content-type, cache token.
|
||||
|
||||
**Not logged:** IP, cookies, Authorization, query string, POST body, user id, session.
|
||||
|
||||
---
|
||||
|
||||
## 7. Expected log volume
|
||||
|
||||
Combined Skinbase access log: ~17 MB so far today, ~31 MB yesterday uncompressed; rotated daily 14 days.
|
||||
|
||||
JSON lines ~300–400 B. Static already excluded. Estimate **~20–40 MB/day** uncompressed, **~300–600 MB** retained (14d) before gzip of older copies (~2–4 MB/day gz). Acceptable; no field trim required.
|
||||
|
||||
Permissions: nginx `0640 www-data adm`. Do not chmod 777. Analyzer may need `sudo`.
|
||||
|
||||
---
|
||||
|
||||
## 8–9. Laravel slow-request logger + threshold
|
||||
|
||||
- Middleware: `LogSlowHttpRequest` (global prepend; **no-op when disabled**)
|
||||
- Channel: `slow-http` → `storage/logs/slow-http.log` daily, **14 days**
|
||||
- Fields: time, method, route_name, route_uri (**template**), status, duration_ms, peak_memory_mb, authenticated bool, inertia bool
|
||||
- **Threshold: 750 ms**
|
||||
|
||||
Rationale: homepage Laravel ~100 ms; guest TTFB 400–500 ms; Academy still 0.6–1.1 s. 750 ms captures Academy/profile/SSR outliers without logging every homepage. If volume is high after 24 h, raise to 1000 ms — do not sample first.
|
||||
|
||||
Disable: `HTTP_SLOW_REQUEST_LOG_ENABLED=false` (default in `.env.example`).
|
||||
|
||||
---
|
||||
|
||||
## 10. SSR timing
|
||||
|
||||
`TimedInertiaSsrGateway` wraps Inertia `HttpGateway` (no vendor patch, SSR not disabled).
|
||||
|
||||
When SSR runs, Server-Timing becomes:
|
||||
|
||||
```
|
||||
app;desc="Laravel";dur=...
|
||||
ssr;desc="Inertia SSR";dur=...
|
||||
```
|
||||
|
||||
`app` includes SSR (nested). Blade home has no `ssr` metric.
|
||||
|
||||
---
|
||||
|
||||
## 11. FPM slowlog
|
||||
|
||||
**NO CHANGE** (already 10s). Optional later: 2s in a **separate** change, not bundled with nginx JSON.
|
||||
|
||||
---
|
||||
|
||||
## 12–14. Analyzer, URI rules, percentiles
|
||||
|
||||
`scripts/analyze-http-performance.php` + `HttpPerformanceLogAnalyzer`
|
||||
|
||||
**Host filter (M12.1):** `--host=skinbase.org` exact match only. Without `--host`, rows aggregate on `host + normalized URI` (hosts are never merged).
|
||||
|
||||
**URI normalization (analysis only):** numeric segments → `{id}`; `/@name` → `/@{user}`; UUID → `{uuid}`; long hex → `{hash}`; `/art/{id}/{slug}` vs `/art/{id}/similar`.
|
||||
|
||||
**Percentile:** nearest-rank, rank = `ceil(p * n)`, 1-indexed. Not an average.
|
||||
|
||||
**Upstream multi-value:** sum numeric comma-separated parts (`0.100, 0.220` → 0.32). `request_time` is authoritative for request latency.
|
||||
|
||||
**499:** counted separately; not treated as 5xx.
|
||||
|
||||
---
|
||||
|
||||
## 15–16. Tests
|
||||
|
||||
```text
|
||||
pest tests/Unit/Http/HttpPerformanceLogAnalyzerTest.php PASS (4)
|
||||
pest tests/Feature/Http/SlowHttpRequestLoggingTest.php PASS (5)
|
||||
```
|
||||
|
||||
Fixtures cover PHP, static `-`, multi-upstream, 200/302/404/499/500, malformed JSON, `/download/artwork/{id}`, `/@{user}`.
|
||||
|
||||
---
|
||||
|
||||
## 17. Files changed
|
||||
|
||||
- `config/http_observability.php`
|
||||
- `config/logging.php`
|
||||
- `.env.example`
|
||||
- `phpunit.xml`
|
||||
- `bootstrap/app.php`
|
||||
- `app/Providers/AppServiceProvider.php`
|
||||
- `app/Http/Middleware/AddServerTiming.php`
|
||||
- `app/Http/Middleware/LogSlowHttpRequest.php`
|
||||
- `app/Support/Http/HttpUriNormalizer.php`
|
||||
- `app/Support/Http/HttpPerformanceLogAnalyzer.php`
|
||||
- `app/Support/Http/TimedInertiaSsrGateway.php`
|
||||
- `scripts/analyze-http-performance.php`
|
||||
- `deploy/nginx/conf.d/skinbase-perf-log-format.conf`
|
||||
- `deploy/nginx/skinbase-vhost-performance-log.snippet`
|
||||
- tests + fixtures
|
||||
- this document
|
||||
|
||||
---
|
||||
|
||||
## 18. Nginx files prepared (not installed)
|
||||
|
||||
See `deploy/nginx/conf.d/skinbase-perf-log-format.conf` and `skinbase-vhost-performance-log.snippet`.
|
||||
|
||||
---
|
||||
|
||||
## 19. Env
|
||||
|
||||
Production recommended:
|
||||
|
||||
```
|
||||
HTTP_SLOW_REQUEST_LOG_ENABLED=true
|
||||
HTTP_SLOW_REQUEST_MS=750
|
||||
HTTP_SLOW_REQUEST_LOG_DAYS=14
|
||||
```
|
||||
|
||||
Local default: **false**.
|
||||
|
||||
---
|
||||
|
||||
## 20. Application deploy (A)
|
||||
|
||||
1. `bash ./sync.sh` (do not `git pull` on production).
|
||||
2. Add the three env vars to production `.env`.
|
||||
3. Config refresh as **skinbase**, only if the host uses config cache:
|
||||
|
||||
`sudo -u skinbase php artisan config:cache`
|
||||
|
||||
4. Do not run artisan as root. Do not `cache:clear` globally unless required.
|
||||
5. PHP-FPM reload only if the deploy process already does (code opcache). No Horizon/Redis/MySQL restart.
|
||||
|
||||
---
|
||||
|
||||
## 21. Nginx deploy (B) — separate from app
|
||||
|
||||
1. Backup: copy `nginx.conf` and `sites-enabled/skinbase.org.conf`.
|
||||
2. Install `skinbase-perf-log-format.conf` into `/etc/nginx/conf.d/`.
|
||||
3. Add **one** extra `access_log` line in the HTTPS server (keep combined).
|
||||
4. `sudo nginx -t` — **stop if fail**.
|
||||
5. `sudo systemctl reload nginx` (**not restart**).
|
||||
6. One `curl` to `/`.
|
||||
7. Confirm a JSON line in `skinbase-performance.log` and a combined line in `skinbase.org-access.log`.
|
||||
8. Confirm no query string / no IP in JSON.
|
||||
9. `sudo -u skinbase php scripts/analyze-http-performance.php --file=/var/log/nginx/skinbase-performance.log --since=1h --top=10`
|
||||
(may need sudo to read `0640 www-data adm`).
|
||||
|
||||
---
|
||||
|
||||
## 22. Optional FPM (C)
|
||||
|
||||
**Skip.** Already logging stacks at 10s.
|
||||
|
||||
---
|
||||
|
||||
## 23. Immediate verification
|
||||
|
||||
```bash
|
||||
curl -sS -D - -o /dev/null --max-time 20 -H 'Accept-Encoding: br' https://skinbase.org/ \
|
||||
| grep -iE 'HTTP/|server-timing|cache-control'
|
||||
|
||||
sudo tail -n 1 /var/log/nginx/skinbase-performance.log
|
||||
sudo tail -n 1 /var/log/nginx/skinbase.org-access.log
|
||||
sudo -u skinbase tail -n 1 /opt/www/virtual/SkinbaseNova/storage/logs/slow-http-$(date +%F).log
|
||||
```
|
||||
|
||||
Expect `Server-Timing` with `app;dur=` (and `ssr;dur=` on Inertia pages). Home may be **below** 750 ms and **not** appear in slow-http.
|
||||
|
||||
---
|
||||
|
||||
## 24. 24h analysis
|
||||
|
||||
```bash
|
||||
sudo php /opt/www/virtual/SkinbaseNova/scripts/analyze-http-performance.php \
|
||||
--file=/var/log/nginx/skinbase-performance.log --since=24h --top=20
|
||||
```
|
||||
|
||||
Then `--status=5xx`, `--status=499`, `--uri=/academy`, `--slow-ms=1000`.
|
||||
|
||||
Do **not** start M13 optimizations from a handful of curls.
|
||||
|
||||
---
|
||||
|
||||
## 25. Rollback
|
||||
|
||||
- App: `HTTP_SLOW_REQUEST_LOG_ENABLED=false` then config cache as skinbase. No code revert required for logging.
|
||||
- Nginx: remove the extra `access_log` line and/or conf.d file, `nginx -t`, `reload`. Combined log untouched.
|
||||
- SSR timing: revert Gateway bind only if a defect appears (SSR behavior unchanged).
|
||||
|
||||
---
|
||||
|
||||
## 26. Risks
|
||||
|
||||
- JSON `access_log` disk (~40 MB/day). Rotation already exists.
|
||||
- Slow-http volume if many pages >750 ms (raise threshold).
|
||||
- `app;dur` will **increase** vs M11’s ~110 ms because timing now includes earlier middleware (this is more accurate, not a regression).
|
||||
- Cloudflare still hides client RTT.
|
||||
|
||||
---
|
||||
|
||||
## 27. NO CHANGE
|
||||
|
||||
FPM `pm.max_children`, Horizon, MySQL, Redis, snapshot indexes, SSR on/off, Telescope, Debugbar, SQL/Redis tracing, forum/collections queues, load tests, nginx restart, public status pages.
|
||||
@@ -0,0 +1,27 @@
|
||||
# M12.2 — Vector Gateway Search Reliability & Circuit Breaker Hotfix
|
||||
|
||||
```text
|
||||
STATUS: COMPLETE (application-only)
|
||||
PRODUCTION: no nginx / env / FPM / Redis / Horizon changes required
|
||||
```
|
||||
|
||||
## Root cause
|
||||
|
||||
`VectorGatewayClient::searchByFileContents()` tripped the global circuit on **any** HTTP failure or thrown exception **before** `AiArtworkVectorSearchService` could run `searchByUrl()`. The fallback then hit `guardCircuit()` and failed immediately. Intermittent 502s (~0.23–0.55s) match fail-fast circuit-open, not cold 1.5–3s searches.
|
||||
|
||||
4xx from the file endpoint also opened the circuit.
|
||||
|
||||
## Circuit policy
|
||||
|
||||
- Combined similar-ai starts only if the circuit is **closed**.
|
||||
- File search, then URL fallback. **Neither method trips or guards the circuit.**
|
||||
- After **both** fail, trip only if **every** gateway failure is circuit-worthy.
|
||||
- Successful fallback **leaves the circuit closed**.
|
||||
- TTL remains **30s**. Search timeouts remain 2s/6s, retries 0.
|
||||
|
||||
Circuit-worthy: connection/timeout, HTTP 408/429/5xx.
|
||||
Not: 400/401/403/404/409/422 and other input 4xx.
|
||||
|
||||
## Deploy
|
||||
|
||||
`bash ./sync.sh` only. Do not reload nginx. Do not clear the circuit key.
|
||||
@@ -0,0 +1,12 @@
|
||||
# M12.3 — Vector Gateway Circuit Trip Attribution
|
||||
|
||||
```text
|
||||
STATUS: COMPLETE (application-only)
|
||||
BEHAVIOR: circuit policy unchanged
|
||||
```
|
||||
|
||||
`tripCircuit()` logs one WARNING only on **closed → open**.
|
||||
|
||||
Message: `Vector gateway circuit opened`
|
||||
|
||||
Victims of an already-open circuit still only get `Vector similarity search failed` with `failure_stage=circuit`.
|
||||
@@ -0,0 +1,7 @@
|
||||
# M12.4 — Require Dual Gateway Failure Before Similar-AI Circuit Trip
|
||||
|
||||
```text
|
||||
STATUS: COMPLETE (application-only)
|
||||
```
|
||||
|
||||
Similar-ai trips the global circuit only when **search_file was actually called** and **both** `search_file` and `search_url` failed circuit-worthy. Isolated URL 502 after a skipped/failed image download no longer opens the 30s breaker.
|
||||
@@ -0,0 +1,79 @@
|
||||
# M2 — Mail Queue Consumption
|
||||
|
||||
## Root cause
|
||||
|
||||
Application mail jobs are pushed to Redis queue **`mail`**. Production Horizon supervisors only consume:
|
||||
|
||||
```text
|
||||
search, default
|
||||
broadcasts, notifications
|
||||
```
|
||||
|
||||
There is **no** `queue:work` process for `mail`. The three items have `attempts=0` — they were never reserved.
|
||||
|
||||
Isolation is **intentional** in code (`onQueue('mail')`) and in `docs/QUEUE.md`. Horizon simply never subscribed.
|
||||
|
||||
## The three queued items (no PII)
|
||||
|
||||
Inspected with `LRANGE`/`LINDEX` via Laravel Redis; only class metadata:
|
||||
|
||||
| # | displayName | maxTries | attempts |
|
||||
| - | ----------- | -------: | -------: |
|
||||
| 0 | `App\Jobs\SendVerificationEmailJob` | 5 | 0 |
|
||||
| 1 | `App\Mail\EmailChangedSecurityAlertMail` | 3 | 0 |
|
||||
| 2 | `App\Mail\EmailChangeVerificationCodeMail` | 3 | 0 |
|
||||
|
||||
`pushedAt` unix times: 1777665006, 1777697989, 1784800925 (older registration/email-change mail sitting until a worker exists).
|
||||
|
||||
## Local vs production
|
||||
|
||||
Same gap on both: Horizon defaults omitted `mail`. Production `config('horizon.defaults.supervisor-*.queue')` confirmed `["search","default"]` and `["broadcasts","notifications"]`.
|
||||
|
||||
Legacy `deploy/supervisor/skinbase-queue.conf` **does** include `mail`, but production runs Horizon, not that unit.
|
||||
|
||||
## Fix
|
||||
|
||||
Added Horizon `supervisor-mail`:
|
||||
|
||||
```text
|
||||
queue: mail
|
||||
connection: redis
|
||||
tries: 5 (matches SendVerificationEmailJob; rate-limit uses release())
|
||||
timeout: 90
|
||||
maxProcesses: 1 local / 2 production
|
||||
```
|
||||
|
||||
Kept isolation so SMTP does not block rec/search workers and so `--tries=1` on default workers cannot fail mail jobs on `release()`.
|
||||
|
||||
## Files changed
|
||||
|
||||
```text
|
||||
config/horizon.php
|
||||
tests/Unit/HorizonMailQueueTest.php
|
||||
docs/QUEUE.md
|
||||
docs/optimization-m2-mail-queue.md
|
||||
```
|
||||
|
||||
## Deployment
|
||||
|
||||
1. Deploy release with `config/horizon.php`.
|
||||
2. `php artisan config:cache` / `optimize` as usual.
|
||||
3. Restart Horizon (Supervisor `skinbase-horizon` / `queue:restart`).
|
||||
4. Confirm `php artisan horizon:status` and a worker with `--queue=mail`.
|
||||
|
||||
No `.env` change required. No Redis delete.
|
||||
|
||||
## Should the 3 messages be processed after deploy?
|
||||
|
||||
**Yes, automatically** once `supervisor-mail` starts — do not delete the list.
|
||||
|
||||
Risks:
|
||||
|
||||
- Recipients may receive **late** verification/security emails.
|
||||
- Tokens/codes may already be expired; user sees a failed verify, not a security hole from us sending mail.
|
||||
- `SendVerificationEmailJob` may still send `RegistrationVerificationMail` (also `mail` queue) — second job is fine once workers run.
|
||||
- Do **not** manually `queue:retry` or `lpop` before Horizon is up.
|
||||
|
||||
## Tests
|
||||
|
||||
`php artisan test --filter=HorizonMailQueue`
|
||||
@@ -0,0 +1,127 @@
|
||||
# M3 — Sitemap Production Audit & Reliability
|
||||
|
||||
## Current architecture
|
||||
|
||||
Routes:
|
||||
|
||||
```text
|
||||
GET /sitemap.xml SitemapController@index
|
||||
GET /sitemaps/{name}.xml SitemapController@show
|
||||
GET /robots.txt RobotsTxtController (Sitemap: {APP_URL}/sitemap.xml)
|
||||
```
|
||||
|
||||
Generation: `php artisan skinbase:sitemaps:generate` (scheduler 10:30 and 22:30) writes static XML on `sitemaps_public` (`public/`).
|
||||
|
||||
Also: `sitemaps:publish`, `validate`, release artifacts, optional live-build fallback.
|
||||
|
||||
Families (no groups/years/videos sitemaps):
|
||||
|
||||
artworks (sharded 10k), academy-*, users (sharded 10k), tags (now shardable at 25k), categories, collections, cards, stories, web-stories, news, news-google, forum-*, static-pages.
|
||||
|
||||
Canonical artwork URLs: `route('art.show')` `/art/{id}/{slug}`. Public/published only.
|
||||
|
||||
## Exact production issue
|
||||
|
||||
`/sitemap.xml` is a **stale file from 2026-05-13**, owned by **klevze** (`-rw-r--r--`). Children (artworks shards, academy, users, …) refresh daily as **skinbase**. Scheduler cannot overwrite the index (`filesystems.sitemaps_public.throw=false`, generate ignored failed `put`). Nginx `try_files $uri` serves the stale root file and **never** hits Laravel. Academy families exist on disk but are **absent from the May index**.
|
||||
|
||||
`location @php` is missing in the live vhost (documented in M0).
|
||||
|
||||
## Local vs production
|
||||
|
||||
Generate command already writes `sitemaps/sitemap.xml` on both. Production **file ownership** plus **nginx `$uri` = public/sitemap.xml** is the operational failure. Local did not fail writes or serve a dual path.
|
||||
|
||||
## Root cause
|
||||
|
||||
1. Index file not writable by `skinbase`.
|
||||
2. `put()` failures ignored.
|
||||
3. Nginx prefers the stale `/sitemap.xml` over `/sitemaps/index.xml` (which did not exist).
|
||||
|
||||
## Implementation
|
||||
|
||||
- Generate writes `sitemaps/index.xml` (new, creatable), plus tries `sitemaps/sitemap.xml` and `sitemap.xml`; **fails the command if none of the index paths write**.
|
||||
- Child writes report failure instead of counting a silent skip as success.
|
||||
- `SitemapController::index()` prefers static `sitemaps/index.xml`.
|
||||
- `deploy/nginx/sitemaps.conf`: `try_files /sitemaps/index.xml /sitemaps/sitemap.xml /index.php?$query_string;` (no `@php`)
|
||||
- Tags builder is shardable (25k) so it cannot exceed 50k URLs without an index.
|
||||
|
||||
## Caching
|
||||
|
||||
Unchanged: nginx `max-age=21600`; Laravel file response uses `sitemaps.cache_ttl_seconds` (900). No HTML cache. Generate is the refresh.
|
||||
|
||||
## Chunking
|
||||
|
||||
Artworks/users/forum-threads/collections/cards/stories already sharded at 10k. Tags 25k. Protocol: 50k URLs / 50MB. Production files are all well under 50MB (largest artworks shard ~3.5MB, tags ~2MB).
|
||||
|
||||
## Deployment
|
||||
|
||||
Production vhost (`/etc/nginx/sites-enabled/skinbase.org.conf`) has **no** `location @php`. PHP-FPM is:
|
||||
|
||||
```nginx
|
||||
location / {
|
||||
try_files $uri $uri/ /index.php?$query_string;
|
||||
}
|
||||
location = /index.php {
|
||||
fastcgi_pass unix:/run/php/php8.4-fpm-skinbase.sock;
|
||||
...
|
||||
}
|
||||
```
|
||||
|
||||
**Do not** use `try_files … @php` (undefined named location). Use `/index.php?$query_string` like `location /`.
|
||||
|
||||
Exact replacement for the two sitemap blocks (currently ~lines 91–110):
|
||||
|
||||
```nginx
|
||||
location = /sitemap.xml {
|
||||
try_files /sitemaps/index.xml /sitemaps/sitemap.xml /index.php?$query_string;
|
||||
|
||||
add_header Cache-Control "public, max-age=21600" always;
|
||||
add_header Content-Type "application/xml; charset=UTF-8" always;
|
||||
etag on;
|
||||
}
|
||||
|
||||
location ~ ^/sitemaps/[A-Za-z0-9_\-]++\.xml$ {
|
||||
try_files $uri /index.php?$query_string;
|
||||
|
||||
add_header Cache-Control "public, max-age=21600" always;
|
||||
add_header Content-Type "application/xml; charset=UTF-8" always;
|
||||
etag on;
|
||||
}
|
||||
```
|
||||
|
||||
Do **not** add `$uri` to the `/sitemap.xml` try_files list: `$uri` is `/sitemap.xml`, the stale klevze-owned file, and would shadow PHP if `index.xml` is missing.
|
||||
|
||||
Order:
|
||||
|
||||
1. Deploy app.
|
||||
2. Run `php artisan skinbase:sitemaps:generate` as `skinbase` so `sitemaps/index.xml` exists.
|
||||
3. Patch the vhost, then:
|
||||
|
||||
```bash
|
||||
sudo nginx -t && sudo systemctl reload nginx
|
||||
```
|
||||
|
||||
`nginx -t` should pass: last `try_files` argument matches the existing front-controller pattern. Introducing `@php` would not.
|
||||
|
||||
Optional: `chown skinbase:skinbase` the old `sitemap.xml`.
|
||||
|
||||
## Legacy index files after `sitemaps/index.xml`
|
||||
|
||||
| Path | Role after M3 |
|
||||
| ---- | ------------- |
|
||||
| `public/sitemaps/index.xml` | **Canonical.** Scheduler can create this even when the May 2026 file is unwritable. Nginx prefers it. |
|
||||
| `public/sitemaps/sitemap.xml` | **Legacy fallback** for `/sitemap.xml` if `index.xml` is missing. Same inode as root `sitemap.xml` on current prod (symlink). Best-effort write; failure must not fail generate if `index.xml` succeeded. |
|
||||
| `public/sitemap.xml` | **Legacy `$uri`.** On prod this is a symlink to the unwritable klevze file. **Not required** once nginx prefers `index.xml`. Generate may skip it. |
|
||||
|
||||
Generate still *attempts* all three; success requires **at least one** index write (`index.xml` is enough).
|
||||
|
||||
## robots.txt
|
||||
|
||||
Already declares `Sitemap: https://skinbase.org/sitemap.xml`. **No robots change** for M3.
|
||||
|
||||
## URLs that should not be in sitemaps
|
||||
|
||||
Already excluded: private/unapproved artworks, inactive users, inactive tags, academy pricing query variants, `/pages/about` duplicate, forbidden `/admin` `/cp` etc. No extra removals in M3.
|
||||
|
||||
## Tests
|
||||
|
||||
`php artisan test tests/Feature/SitemapTest.php` (run after this doc).
|
||||
@@ -0,0 +1,178 @@
|
||||
# M4 — Recommendation Snapshot & Read Path
|
||||
|
||||
**Snapshot table needed: NO**
|
||||
|
||||
Investigation only plus a regression test. No production changes. No new tables.
|
||||
|
||||
---
|
||||
|
||||
## 1. What “snapshot table” meant
|
||||
|
||||
Two different things were previously named “snapshot”:
|
||||
|
||||
| Table | Size (prod) | Role |
|
||||
| ----- | ----------: | ---- |
|
||||
| `artwork_metric_snapshots_hourly` | **1855 MB / 8.6M rows** | Heat/ranking hourly totals (M0 DB residual). **Not** similar-art serving. |
|
||||
| `rec_artwork_recs` | **11.6 MB / 27.7k rows** | Precomputed similar lists (JSON `recs`) read by the API. |
|
||||
|
||||
M4 is about the **recommendation read path**, not heat snapshots. A second copy of `rec_artwork_recs` would not address the 1.85 GB hourly table. That remains a later ranking/ops milestone.
|
||||
|
||||
---
|
||||
|
||||
## 2. Current data model
|
||||
|
||||
```text
|
||||
rec_item_pairs
|
||||
(a_artwork_id, b_artwork_id) UNIQUE
|
||||
weight, updated_at
|
||||
~74k rows, 16 MB
|
||||
Written by RecBuildItemPairsFromFavouritesJob (DELETE all, then upsert chunks)
|
||||
|
||||
rec_artwork_recs
|
||||
UNIQUE (artwork_id, rec_type, model_version)
|
||||
recs JSON (ordered IDs), computed_at
|
||||
rec_type: similar_tags | similar_behavior | similar_hybrid | similar_visual
|
||||
Written by RecComputeSimilar* via updateOrCreate (one row at a time)
|
||||
|
||||
user_recommendation_cache
|
||||
personalized “For You” blobs (~67 rows, 2.5 MB) — separate from similar-art
|
||||
```
|
||||
|
||||
Frontend/API similar-art:
|
||||
|
||||
```text
|
||||
GET /api/art/{id}/similar
|
||||
GET /art/{id}/similar
|
||||
→ HybridSimilarArtworksService::forArtwork
|
||||
Cache::remember rec:artwork:{id}:similar:{model} TTL 6h
|
||||
SELECT rec_artwork_recs WHERE artwork_id + rec_type + model_version (const)
|
||||
hydrate Artwork whereIn(ids) public+published
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 3. Production row counts (2026-08-23, read-only)
|
||||
|
||||
| rec_type | rows | computed_at span |
|
||||
| -------- | ---: | ---------------- |
|
||||
| similar_hybrid | 13902 | 2026-03-15 → 2026-08-23 14:36 |
|
||||
| similar_tags | 9479 | 2026-03-15 → 2026-08-23 14:58 |
|
||||
| similar_behavior | 4851 | 2026-04-21 → 2026-08-23 14:34 |
|
||||
|
||||
~50k public artworks exist; lists are **incomplete** because nightly catalog jobs were dying (M1), not because of missing snapshots. Coverage will grow after M1 deploy.
|
||||
|
||||
---
|
||||
|
||||
## 4. Read/write flow
|
||||
|
||||
**Read:** unique const lookup (1 row) + Redis/cache 6h + one `whereIn` hydrate. Cheap.
|
||||
|
||||
**Write (M1 batches):** 200 artworks/job, `updateOrCreate` per artwork. No truncate of `rec_artwork_recs`. Unique key replaces one JSON list atomically.
|
||||
|
||||
**Write (pairs):** `DELETE FROM rec_item_pairs` then rebuild every 4 hours. Readers of pairs are **rebuild jobs**, not the HTTP similar endpoint.
|
||||
|
||||
**Partial visibility:** during a catalog rebuild, some artworks have today’s `computed_at`, others yesterday. HTTP cache may keep old IDs for 6h. There is no empty-table window for `rec_artwork_recs`. Mixing generations is **per artwork**, not a torn JSON array.
|
||||
|
||||
**Transactions:** none wrapping the full catalog. Each `updateOrCreate` is its own statement.
|
||||
|
||||
---
|
||||
|
||||
## 5. Query plans (production EXPLAIN)
|
||||
|
||||
Read:
|
||||
|
||||
```text
|
||||
type=const
|
||||
key=rec_artwork_recs_artwork_id_rec_type_model_version_unique
|
||||
rows=1
|
||||
```
|
||||
|
||||
Pairs UNION used by behavior rebuild (not HTTP):
|
||||
|
||||
```text
|
||||
ref on unique (a_artwork_id, b_artwork_id)
|
||||
ref on b_artwork_id index
|
||||
UNION RESULT: Using temporary; Using filesort (tiny LIMIT 90)
|
||||
```
|
||||
|
||||
No full table scan on the similar-art HTTP path.
|
||||
|
||||
---
|
||||
|
||||
## 6. Bottleneck (exact)
|
||||
|
||||
| Concern | Present? |
|
||||
| ------- | -------- |
|
||||
| Expensive HTTP reads | **No** — const unique + cache |
|
||||
| Partial rec JSON during rebuild | **No** — row replace |
|
||||
| Mixed old/new *across* artworks during overnight batch | **Yes, benign** with 6h cache |
|
||||
| delete/reinsert on `rec_artwork_recs` | **No** |
|
||||
| delete/reinsert on `rec_item_pairs` | **Yes**, 74k rows / 4h; not on the HTTP path |
|
||||
| Poor indexing | **No** on read |
|
||||
| Excessive Redis | One key per artwork, 6h TTL — normal |
|
||||
| Dual generations mixed in one response | **No** |
|
||||
|
||||
**Exact bottleneck for similar-art quality:** incomplete `rec_artwork_recs` coverage from failed nightly jobs (M1), not read-path schema.
|
||||
|
||||
**Separate large table:** `artwork_metric_snapshots_hourly` 1.85 GB — ranking/heat, **out of M4 schema scope**.
|
||||
|
||||
---
|
||||
|
||||
## 7. Snapshot table?
|
||||
|
||||
```text
|
||||
YES/NO: NO
|
||||
```
|
||||
|
||||
Would double a **12 MB** table, add cutover complexity, and not fix M1 coverage or the 1.85 GB heat snapshots.
|
||||
|
||||
### Simpler alternatives (do not implement in M4 unless needed later)
|
||||
|
||||
1. **Deploy M1** so batches complete; coverage grows via `updateOrCreate`.
|
||||
2. If `rec_item_pairs` DELETE causes empty-pair windows during overlap with behavior jobs: replace global delete with rebuild-into-temp + `RENAME TABLE` **on that 16 MB table only** — still not a rec snapshot table.
|
||||
3. Heat snapshots: partition/prune `artwork_metric_snapshots_hourly` in a dedicated milestone.
|
||||
|
||||
---
|
||||
|
||||
## 8. Proposed schema if YES
|
||||
|
||||
N/A.
|
||||
|
||||
---
|
||||
|
||||
## 9–12. Migration / cutover / cleanup / deploy
|
||||
|
||||
N/A for snapshot tables. M1 deploy remains the rec-job fix.
|
||||
|
||||
Rollback: none (no schema).
|
||||
|
||||
---
|
||||
|
||||
## 13. Tests
|
||||
|
||||
Added: `keeps other rec types and other artworks visible when one row is rebuilt` in `tests/Feature/Recommendations/SimilarArtworksHybridTest.php`.
|
||||
|
||||
Existing hybrid tests already cover fallback chain, order, author cap, API `type=`.
|
||||
|
||||
---
|
||||
|
||||
## 14. Files changed
|
||||
|
||||
```text
|
||||
tests/Feature/Recommendations/SimilarArtworksHybridTest.php
|
||||
docs/optimization-m4-recommendation-snapshots.md
|
||||
```
|
||||
|
||||
No migrations, no rec model changes.
|
||||
|
||||
---
|
||||
|
||||
## 15. Risks
|
||||
|
||||
None from schema (none added). Residual: pairs `DELETE` vs behavior schedule overlap; 6h cache lag after rebuild; heat snapshot table still large.
|
||||
|
||||
---
|
||||
|
||||
## Production
|
||||
|
||||
Read-only EXPLAIN + `information_schema` + grouped counts. Temp `/tmp` script removed. No deploys, no writes to rec tables.
|
||||
@@ -0,0 +1,329 @@
|
||||
# M5 — Artwork Metric Snapshot Retention & Growth
|
||||
|
||||
```text
|
||||
STATUS: COMPLETE (local implementation; not deployed)
|
||||
PRODUCTION WRITES: none
|
||||
PRODUCTION DELETES: none
|
||||
OPTIMIZE TABLE: not run
|
||||
PARTITIONING: not added
|
||||
NEW ANALYTICS TABLE: not created
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Verdict
|
||||
|
||||
```text
|
||||
Pruning needed: YES
|
||||
Recommended retention: 30 days
|
||||
Downsampling needed: NO
|
||||
Partitioning needed: NO
|
||||
Indexes to add: none
|
||||
Indexes to remove: optional later (redundant idx_artwork_bucket)
|
||||
Expected storage vs unlimited: large reduction
|
||||
Expected storage vs current 7d prune: INCREASE (~4.3×) if retention is raised to 30d
|
||||
```
|
||||
|
||||
Hourly snapshots are already pruned daily with `--keep-days=7` (unbounded single `DELETE`). That window is **too short** for monthly leaderboards and Studio 30d views. M5 makes prune batched + configurable and sets default retention to **30 days**.
|
||||
|
||||
---
|
||||
|
||||
## 1. Table schema (production)
|
||||
|
||||
```sql
|
||||
CREATE TABLE `artwork_metric_snapshots_hourly` (
|
||||
`id` bigint unsigned NOT NULL AUTO_INCREMENT,
|
||||
`artwork_id` bigint unsigned NOT NULL,
|
||||
`bucket_hour` datetime NOT NULL,
|
||||
`views_count` bigint unsigned NOT NULL DEFAULT '0',
|
||||
`downloads_count` bigint unsigned NOT NULL DEFAULT '0',
|
||||
`favourites_count` bigint unsigned NOT NULL DEFAULT '0',
|
||||
`comments_count` bigint unsigned NOT NULL DEFAULT '0',
|
||||
`shares_count` bigint unsigned NOT NULL DEFAULT '0',
|
||||
`created_at` timestamp NOT NULL DEFAULT CURRENT_TIMESTAMP,
|
||||
PRIMARY KEY (`id`),
|
||||
UNIQUE KEY `uq_artwork_bucket` (`artwork_id`,`bucket_hour`),
|
||||
KEY `idx_bucket_hour` (`bucket_hour`),
|
||||
KEY `idx_artwork_bucket` (`artwork_id`,`bucket_hour`),
|
||||
KEY `idx_bucket_artwork` (`bucket_hour`,`artwork_id`),
|
||||
CONSTRAINT `artwork_metric_snapshots_hourly_artwork_id_foreign`
|
||||
FOREIGN KEY (`artwork_id`) REFERENCES `artworks` (`id`) ON DELETE CASCADE
|
||||
) ENGINE=InnoDB
|
||||
```
|
||||
|
||||
Not partitioned (`CREATE_OPTIONS` empty; `PARTITIONS.PARTITION_NAME` NULL).
|
||||
|
||||
Duplicates: **none** (unique `(artwork_id, bucket_hour)`). `has_duplicate_groups=0`.
|
||||
|
||||
---
|
||||
|
||||
## 2. Production size & range (read-only, 2026-08-23 15:07 UTC)
|
||||
|
||||
| Metric | Value |
|
||||
| ------ | ----- |
|
||||
| Exact `COUNT(*)` | **8,918,000** |
|
||||
| `information_schema` est. | 8,622,217 |
|
||||
| Data | 740.03 MB |
|
||||
| Index | 1115.47 MB |
|
||||
| Total | **1855.50 MB** |
|
||||
| InnoDB `DATA_FREE` | 121.00 MB |
|
||||
| Distinct artworks | 49,822 |
|
||||
| Distinct hours | 179 |
|
||||
| Oldest `bucket_hour` | 2026-08-16 05:00:00 |
|
||||
| Newest `bucket_hour` | 2026-08-23 15:00:00 |
|
||||
| Rows per hour (stable) | **49,822** |
|
||||
| Rows / day | **1,195,728** |
|
||||
| AUTO_INCREMENT | 153,018,248 (historical insert churn after prune) |
|
||||
|
||||
Rows older than 7 days at query time: 548,031 (~11 extra hours until the next 04:00 prune). Oldest row is 05:00 on the day after the last 04:00 cutoff — **daily prune is running**.
|
||||
|
||||
---
|
||||
|
||||
## 3. Growth estimates
|
||||
|
||||
Assumes current write rate stays ~49,822 rows/hour and size scales linearly (~249 MB/day from 1855.5 MB / 7.45 days).
|
||||
|
||||
| Horizon | Rows | Approx size |
|
||||
| ------- | ----: | ----------: |
|
||||
| Current (~7.5 d) | 8.9M | 1.86 GB |
|
||||
| 24 hours | 1.2M | ~0.25 GB |
|
||||
| 7 days | 8.4M | ~1.75 GB |
|
||||
| **30 days** | **35.9M** | **~7.5 GB** |
|
||||
| 90 days | 107.6M | ~22.4 GB |
|
||||
| 365 days | 436.4M | ~91 GB |
|
||||
| Unlimited 1y of this rate | same as 365d | ~91 GB |
|
||||
|
||||
MB/day ≈ **249 MB/day** (data+index) at current density.
|
||||
|
||||
---
|
||||
|
||||
## 4. Writers (exact)
|
||||
|
||||
| Writer | When | What |
|
||||
| ------ | ---- | ---- |
|
||||
| `nova:metrics-snapshot-hourly` (`MetricsSnapshotHourlyCommand`) | Scheduler `hourlyAt(2)` | Upsert one row per eligible artwork for `now()->startOfHour()` |
|
||||
| Eligibility | `--days=60` default | `artworks.created_at` last 60 days **OR** `artwork_stats.ranking_score > 0`; approved, not deleted |
|
||||
|
||||
Chunk 1000. Unique key makes reruns idempotent. **No other writers** (no jobs, no HTTP, no observers).
|
||||
|
||||
That eligibility (ranking_score>0) explains ~50k IDs/hour, not only last-60-day artworks.
|
||||
|
||||
---
|
||||
|
||||
## 5. Readers (exact)
|
||||
|
||||
| Consumer | Lookback | Notes |
|
||||
| -------- | -------- | ----- |
|
||||
| `RecalculateHeatCommand` (`nova:recalculate-heat` every 15 min) | **24h** | DISTINCT artwork_id + load window; writes `artwork_stats.heat_score` |
|
||||
| `HomepageService::risingRecentActivitySubquery` | **24h** | MAX-MIN deltas |
|
||||
| `DiscoverController` rising | **24h** | same pattern |
|
||||
| `DiscoverFeedController` (RSS) | **24h** | same pattern |
|
||||
| `LeaderboardService` daily/weekly/**monthly** | **1 / 7 / ~30 days** | `MAX-MIN` cumulative counts grouped by artwork; all-time uses `artwork_stats`, not this table |
|
||||
| `StudioMetricsService::getDashboardKpis` | **30 days** | `SUM(views_count)` — **semantically wrong** (views_count is cumulative totals, not hourly increments). Falls back to lifetime stats if 0. |
|
||||
| `CreatorJourneyService::biggestDownloadSpike` | **all retained hours** for the creator’s public artworks | No time filter; bounded only by prune |
|
||||
| `nova:prune-metric-snapshots` | cutoff | writer of deletes |
|
||||
|
||||
Frontend/API do not query this table directly; they read heat/leaderboard/studio payloads derived from it.
|
||||
|
||||
No daily/weekly/monthly artwork metric aggregate table exists (only collection daily stats elsewhere).
|
||||
|
||||
---
|
||||
|
||||
## 6. Required historical retention (not assumed)
|
||||
|
||||
| Need | Depth |
|
||||
| ---- | ----- |
|
||||
| Heat / rising homepage / discover / RSS | 24 hours |
|
||||
| Weekly leaderboards | 7 days |
|
||||
| Monthly leaderboards | **~30 days** |
|
||||
| Studio “views 30d” query | 30 days (even though the SUM is a bad formula) |
|
||||
| Journey download spike | prefers longer; not a product requirement for unlimited |
|
||||
| Rec similar-art serving | **none** (M4) |
|
||||
|
||||
**Required: 30 days** of hourly rows if monthly leaderboards and Studio 30d stay on this table.
|
||||
|
||||
Unlimited / 1 year / 90 days: **not required** by application code.
|
||||
|
||||
Current production `--keep-days=7` **under-retains** monthly leaderboard snapshot deltas.
|
||||
|
||||
---
|
||||
|
||||
## 7. Production query plans (EXPLAIN)
|
||||
|
||||
Heat DISTINCT 24h:
|
||||
|
||||
```text
|
||||
type=range key=idx_bucket_artwork rows~2.46M
|
||||
Using where; Using index; Using temporary
|
||||
```
|
||||
|
||||
Rising GROUP BY 24h:
|
||||
|
||||
```text
|
||||
type=range key=idx_bucket_artwork rows~2.46M
|
||||
Using index condition; Using temporary
|
||||
```
|
||||
|
||||
Studio 30d for `user_id=1`:
|
||||
|
||||
```text
|
||||
artworks: ref artworks_user_id_index (~449 rows)
|
||||
snapshots: ref uq_artwork_bucket (artwork_id) ~411 rows, Using index condition
|
||||
```
|
||||
|
||||
Leaderboard monthly GROUP BY (full table, 30d > retained 7d so it scans retained set):
|
||||
|
||||
```text
|
||||
type=index key=uq_artwork_bucket rows~8.87M Using where
|
||||
```
|
||||
|
||||
Prune count `bucket_hour < now()-7d`:
|
||||
|
||||
```text
|
||||
type=range key=idx_bucket_artwork rows~1.1M Using where; Using index
|
||||
```
|
||||
|
||||
Indexes are used. 24h range still estimates millions of rows because ~50k artworks × 24 hours ≈ 1.2M (optimizer overestimate 2.46M). Acceptable for hourly/15-min jobs; leaderboard monthly on 30d would scan ~36M rows if retention grows — still one scheduled hourly job, not request path.
|
||||
|
||||
---
|
||||
|
||||
## 8. Pruning / downsample / partition / archive
|
||||
|
||||
### Pruning: YES
|
||||
|
||||
Already scheduled. Replace unbounded `DELETE WHERE bucket_hour < cutoff` with PK-id batches.
|
||||
|
||||
### Downsampling: NO
|
||||
|
||||
No existing daily artwork snapshot table. Monthly scores only need MIN/MAX of cumulative counters in the window; hourly rows are convenient. A second table is more operational cost than 30d hourly until size is painful (~7.5 GB still fits RAM buffer pool 10 GB). Revisit if table exceeds ~15–20 GB.
|
||||
|
||||
### Archive: NO
|
||||
|
||||
Nothing reads off-box history.
|
||||
|
||||
### Partitioning: NO
|
||||
|
||||
Rolling 7–30 day window + batched DELETE is simpler than `PARTITION BY RANGE (TO_DAYS(bucket_hour))` + monthly `ALTER`/`DROP PARTITION`. Not justified at 1.86 GB (or 7.5 GB).
|
||||
|
||||
---
|
||||
|
||||
## 9. Delete vs InnoDB free space vs filesystem
|
||||
|
||||
- `DELETE` removes rows; InnoDB keeps pages in the tablespace (`DATA_FREE` today 121 MB).
|
||||
- Disk file (`ibd`) typically **does not shrink**.
|
||||
- Filesystem reclaim needs `OPTIMIZE TABLE` / `ALTER TABLE ... ENGINE=InnoDB` (rebuild) — **do not run automatically on production**.
|
||||
- After raising retention, size grows; after later lowering it, expect logical free space, not a smaller `.ibd` until a rebuild is planned in a maintenance window.
|
||||
|
||||
---
|
||||
|
||||
## 10. Implementation (local)
|
||||
|
||||
- Config `config/metrics.php` + env `ARTWORK_METRIC_HOURLY_RETENTION_DAYS` default **30**.
|
||||
- `PruneMetricSnapshotsCommand`: count eligible → loop `SELECT id … LIMIT chunk` → `DELETE WHERE id IN (…)`, sleep, logs per batch, `--dry-run`, `--max-batches`.
|
||||
- Scheduler: `nova:prune-metric-snapshots` **without** hardcoded `--keep-days=7`.
|
||||
- Tests for keep-days, batches, idempotence, dry-run, config default, max-batches.
|
||||
|
||||
### Production-safe properties
|
||||
|
||||
| Requirement | How |
|
||||
| ----------- | --- |
|
||||
| Bounded batches | `--chunk` default 5000 |
|
||||
| No giant DELETE | one `whereIn` per batch |
|
||||
| Avoid long locks | small PK deletes + `--sleep-ms=50` |
|
||||
| Scheduler-safe | `withoutOverlapping`; daily 04:00 |
|
||||
| Idempotent | re-run deletes 0 extra rows |
|
||||
| Observable | info logs per batch + completed |
|
||||
| Configurable | env + `--keep-days` override |
|
||||
|
||||
---
|
||||
|
||||
## 11. Indexes
|
||||
|
||||
Keep:
|
||||
|
||||
- `PRIMARY (id)` — batched prune
|
||||
- `uq_artwork_bucket (artwork_id, bucket_hour)` — upsert + per-artwork time
|
||||
- `idx_bucket_artwork (bucket_hour, artwork_id)` — heat/rising/prune range
|
||||
|
||||
Optional later (not in this milestone):
|
||||
|
||||
- Drop `idx_artwork_bucket` — duplicate of unique key
|
||||
- Drop `idx_bucket_hour` — prefix of `idx_bucket_artwork`
|
||||
|
||||
Do not add indexes. Do not drop on production in M5 (online DDL on 1.85 GB).
|
||||
|
||||
---
|
||||
|
||||
## 12. Expected storage after M5 deploy
|
||||
|
||||
If `ARTWORK_METRIC_HOURLY_RETENTION_DAYS=30` on production:
|
||||
|
||||
- Table grows from ~1.86 GB toward **~7.5 GB** over ~3 weeks.
|
||||
- That is **correctness**, not reduction.
|
||||
|
||||
If production must stay small, set `ARTWORK_METRIC_HOURLY_RETENTION_DAYS=7` and accept monthly leaderboard snapshot windows of only 7 days. **Do not** leave the scheduler at 7 days while believing Studio/monthly boards have 30d of hourly history.
|
||||
|
||||
Storage reduction vs never pruning: **~91 GB/year avoided**.
|
||||
|
||||
---
|
||||
|
||||
## 13. Deployment (later milestone — not M5)
|
||||
|
||||
1. Ship code + `config/metrics.php`.
|
||||
2. Set env on server: `ARTWORK_METRIC_HOURLY_RETENTION_DAYS=30` (or 7 if size is prioritized).
|
||||
3. `php artisan config:cache` as the deploy already does.
|
||||
4. Do **not** run prune by hand on first deploy unless a dry-run is wanted: `php artisan nova:prune-metric-snapshots --dry-run`.
|
||||
5. Do **not** `OPTIMIZE TABLE`.
|
||||
6. Watch next 04:00 run logs for batch counts.
|
||||
|
||||
If raising 7 → 30: prune deletes almost nothing until the table ages to 30 days.
|
||||
|
||||
---
|
||||
|
||||
## 14. Rollback
|
||||
|
||||
- Restore previous command (single DELETE) and scheduler `--keep-days=7`.
|
||||
- Or set env back to `7` with the new command (preferred).
|
||||
- No schema change to roll back.
|
||||
|
||||
---
|
||||
|
||||
## 15. Tests
|
||||
|
||||
- `tests/Feature/PruneMetricSnapshotsCommandTest.php` (new)
|
||||
- Existing `tests/Feature/RisingEngineTest.php` prune example still valid with `--keep-days=7`
|
||||
|
||||
---
|
||||
|
||||
## 16. Files changed
|
||||
|
||||
```text
|
||||
config/metrics.php (new)
|
||||
app/Console/Commands/PruneMetricSnapshotsCommand.php (batched prune)
|
||||
routes/console.php (no hardcoded 7)
|
||||
.env.example
|
||||
docs/cli-reference.md
|
||||
docs/optimization-m5-metric-snapshot-retention.md (new)
|
||||
tests/Feature/PruneMetricSnapshotsCommandTest.php (new)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 17. Risks
|
||||
|
||||
- Raising retention to 30d **increases** MySQL size and leaderboard monthly scan cost.
|
||||
- Studio 30d `SUM(views_count)` remains an incorrect metric (cumulative summed across hours). Out of M5 scope.
|
||||
- Journey spike still limited to retained hours.
|
||||
- Batched prune of ~1.2M rows/day at chunk 5000 ≈ 240 batches; 50ms sleep ≈ 12s extra plus delete time. Fine for 04:00.
|
||||
- Redundant indexes still cost ~index-heavy 1.1 GB; dropping them is a later DDL decision.
|
||||
|
||||
---
|
||||
|
||||
## 18. What was not done
|
||||
|
||||
- No production SQL writes/deletes
|
||||
- No `OPTIMIZE TABLE`
|
||||
- No partitioning
|
||||
- No new aggregate table
|
||||
- No change to snapshot writer eligibility (`--days=60` + ranking_score)
|
||||
- M1–M4 still undeployed
|
||||
@@ -0,0 +1,195 @@
|
||||
# M5.1 — Metric Snapshot Correctness & 30-Day Readiness
|
||||
|
||||
```text
|
||||
STATUS: COMPLETE (local; not deployed)
|
||||
PRODUCTION WRITES: none
|
||||
OPTIMIZE TABLE: not run
|
||||
PARTITIONING: not added
|
||||
NEW TABLES: none
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Verdict
|
||||
|
||||
```text
|
||||
Safe to change production retention 7 → 30 after this code ships: YES
|
||||
Recommended env: ARTWORK_METRIC_HOURLY_RETENTION_DAYS=30
|
||||
M5 prune batches: verified
|
||||
Hourly writer vs prune overlap: staggered (prune 04:25, snapshot hourlyAt 2)
|
||||
Monthly GROUP BY at ~36M rows: acceptable as a scheduled job; not request-path
|
||||
```
|
||||
|
||||
Production prune is already at 30 days (`eligible=0`); **this KPI/delta code is still local and must ship before treating Studio/monthly numbers as correct 30-day metrics.**
|
||||
|
||||
---
|
||||
|
||||
## 1. Cumulative-counter bugs found
|
||||
|
||||
| Location | Bug | Status |
|
||||
| -------- | --- | ------ |
|
||||
| `StudioMetricsService::getDashboardKpis` | `SUM(views_count)` over cumulative hourly rows | **FIXED** |
|
||||
| Same method | If SUM was 0, **fallback to lifetime** `artwork_stats.views` as `views_30d` | **REMOVED** |
|
||||
| Same method | `favourites_30d` / `shares_30d` were **lifetime** totals | **FIXED** |
|
||||
| `CreatorStudioOverviewService` (live Studio) | 30d KPI keys were **lifetime** module totals | **FIXED** (artwork hourly window + coverage) |
|
||||
| `LeaderboardService` artwork daily/weekly/monthly | `MAX−MIN` **inside** the window only — undercounts **new** artworks (first snapshot already has views) | **FIXED** (latest − pre-window baseline, or 0 if born in period) |
|
||||
| Heat `RecalculateHeatCommand` | Single snapshot without earlier hour → 0 heat | **intentional** (momentum, not 30d KPI) |
|
||||
| Rising homepage/discover/RSS | `MAX−MIN` over 24h | **acceptable** for 24h momentum; first hour of a brand-new work is undercounted |
|
||||
| Nova cards / collections `SUM(views_count)` | Other tables | **not a snapshot bug** |
|
||||
|
||||
No remaining `SUM` of `artwork_metric_snapshots_hourly` cumulative columns.
|
||||
|
||||
---
|
||||
|
||||
## 2. Studio `views_30d` fix
|
||||
|
||||
For each artwork:
|
||||
|
||||
```text
|
||||
latest = MAX(counter) in [start, now]
|
||||
baseline =
|
||||
0 if published_at/created_at >= start
|
||||
else snapshot at/immediately before start (last 36h before start)
|
||||
else MIN(counter) in window -- warm-up / missing hours; NOT lifetime
|
||||
delta = GREATEST(latest - baseline, 0)
|
||||
```
|
||||
|
||||
Then SUM those **deltas** across the creator’s artworks.
|
||||
|
||||
- Created during the period: first snapshot of 40 + later 90 → **90**
|
||||
- Older work with pre-window 200 and latest 250 → **50**
|
||||
- Older work, one in-window snapshot of 400 (no baseline) → **0** (do not report lifetime)
|
||||
- Zero growth stays **0** (no lifetime fallback)
|
||||
|
||||
Coverage is exposed so a 7.5-day table is not labeled as a full 30 days.
|
||||
|
||||
---
|
||||
|
||||
## 3. Warm-up behavior
|
||||
|
||||
Production history today starts **2026-08-16** (~7.5 days). After env=30, history stays incomplete until ~2026-09-15.
|
||||
|
||||
UI/API now:
|
||||
|
||||
- Still compute the delta from **available** hours (do not invent 30d).
|
||||
- Expose `kpis.snapshot_window`:
|
||||
- `requested_days` (30)
|
||||
- `table_coverage_days`
|
||||
- `user_coverage_days`
|
||||
- `window_complete` (true only when table span ≥ 30 days minus one hour)
|
||||
- oldest/newest buckets
|
||||
- Studio dashboard hint: **“Based on 7.5 of 30 days”** until complete; then **“Last 30 days”**.
|
||||
|
||||
Monthly leaderboards already score from `bucket_hour >= now-1 month` (whatever rows exist). They are internally consistent but **short** until warm-up ends. No separate leaderboard badge in this milestone (not a Studio KPI). Do not treat those boards as a full calendar month until coverage is complete.
|
||||
|
||||
---
|
||||
|
||||
## 4. Creator journey
|
||||
|
||||
`biggestDownloadSpike` loaded **all** retained hours for every public artwork (unbounded vs retention). After 30d that is ~720 hours × N artworks.
|
||||
|
||||
**Decision:** explicit lookback = `metrics.hourly_snapshot_retention_days` (default 30).
|
||||
|
||||
- Career “all-time spike” cannot exist beyond retention anyway.
|
||||
- Query size stays bounded when retention is raised later.
|
||||
- Milestone copy now says the retained hourly window.
|
||||
- Fixture snapshots moved to recent hours so the test still sees a spike.
|
||||
|
||||
---
|
||||
|
||||
## 5. Monthly leaderboard at 30-day scale
|
||||
|
||||
M5 EXPLAIN (production, 8.9M rows, 30d predicate covering the whole table):
|
||||
|
||||
```text
|
||||
type=index key=uq_artwork_bucket rows~8.87M Using where
|
||||
GROUP BY artwork_id MAX-MIN
|
||||
```
|
||||
|
||||
At 36M rows the plan stays a covering unique-index scan, ~4× current row volume. `leaderboards:refresh` is hourly, `withoutOverlapping`, not on the HTTP path.
|
||||
|
||||
**Acceptable** for a scheduled job. If wall time exceeds a few minutes after warm-up, rewrite to range on `idx_bucket_artwork` or pre-aggregate — not required now. No load test was run on production.
|
||||
|
||||
24h rising/heat already use `idx_bucket_artwork` range (~1.2M rows) and stay fine.
|
||||
|
||||
---
|
||||
|
||||
## 6. Pruning (verified in code)
|
||||
|
||||
| Requirement | Implementation |
|
||||
| ----------- | ---------------- |
|
||||
| Default 30 days | `config/metrics.php` / `ARTWORK_METRIC_HOURLY_RETENTION_DAYS` |
|
||||
| Bounded deletes | `--chunk` default **5000**, PK `id IN (...)` |
|
||||
| Sleep | `--sleep-ms` default **50** |
|
||||
| No giant transaction | one DELETE per batch |
|
||||
| No self-overlap | `withoutOverlapping(120)` |
|
||||
| Logs | start, each batch, completed |
|
||||
| Resume next day | cutoff is `now - keep-days`; leftover rows delete on the next run |
|
||||
| Scheduler time | **04:25** (was 04:00) |
|
||||
|
||||
---
|
||||
|
||||
## 7. Snapshot writer vs prune
|
||||
|
||||
| Job | Time |
|
||||
| --- | ---- |
|
||||
| `nova:metrics-snapshot-hourly` | `hourlyAt(2)` → **04:02** upserts **current** hour (~50k rows) |
|
||||
| `nova:prune-metric-snapshots` | **04:25** deletes `bucket_hour < now-30d` |
|
||||
|
||||
They no longer share 04:00–04:02. Locks: prune deletes old PK/`bucket_hour` range; writer upserts `(artwork_id, current hour)`. Different unique-key tuples. Residual risk is IO only, not a correctness lock.
|
||||
|
||||
`withoutOverlapping` is per schedule **name**, so the two jobs can still run together if prune overruns into 05:02. Stagger + 50ms sleeps keep that unlikely.
|
||||
|
||||
---
|
||||
|
||||
## 8. Deployment sequence (later; not M5.1)
|
||||
|
||||
1. Deploy this code (window helper, Studio KPIs, journey lookback, prune 04:25).
|
||||
2. Set `ARTWORK_METRIC_HOURLY_RETENTION_DAYS=30`.
|
||||
3. `config:cache`.
|
||||
4. Optional: `php artisan nova:prune-metric-snapshots --dry-run` (read-only).
|
||||
5. Do **not** OPTIMIZE TABLE.
|
||||
6. Expect table growth toward ~7.5 GB over ~3 weeks; Studio shows “N of 30 days” until then.
|
||||
|
||||
**Rollback:** env=7; Studio KPIs still MAX−MIN (correct on whatever history exists).
|
||||
|
||||
---
|
||||
|
||||
## 9. Tests
|
||||
|
||||
- `tests/Feature/Metrics/ArtworkHourlySnapshotWindowTest.php` (SUM vs delta, new artwork, missing baseline, partial history, zero growth)
|
||||
- `tests/Feature/Metrics/LeaderboardSnapshotDeltaTest.php` (daily/weekly/monthly)
|
||||
- `tests/Feature/StudioTest.php` (snapshot_window on dashboard)
|
||||
- creator journey retention window
|
||||
- existing prune tests
|
||||
|
||||
---
|
||||
|
||||
## 10. Files changed
|
||||
|
||||
```text
|
||||
app/Services/Metrics/ArtworkHourlySnapshotWindow.php
|
||||
app/Services/LeaderboardService.php
|
||||
app/Services/Studio/StudioMetricsService.php
|
||||
app/Services/Studio/CreatorStudioOverviewService.php
|
||||
app/Services/Profile/CreatorJourneyService.php
|
||||
resources/js/Pages/Studio/StudioDashboard.jsx
|
||||
routes/console.php
|
||||
config/metrics.php
|
||||
tests/Feature/Metrics/ArtworkHourlySnapshotWindowTest.php
|
||||
tests/Feature/Metrics/LeaderboardSnapshotDeltaTest.php
|
||||
tests/Feature/StudioTest.php
|
||||
tests/Feature/Profile/CreatorJourneyTest.php
|
||||
docs/optimization-m5.1-metric-snapshot-correctness.md
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 11. Risks
|
||||
|
||||
- Studio 30d KPIs now **artwork-only** (hourly snapshots). Cards/collections/stories lifetime totals no longer inflate those tiles.
|
||||
- Cached `studio.kpi.{id}` 5 minutes.
|
||||
- Table-level `window_complete` is global, not per-user (correct for production warm-up).
|
||||
- Monthly boards remain unlabeled for incomplete months.
|
||||
|
||||
M5+M5.1 together are **safe to deploy** with env=30 after this code is on the server. Do not deploy env=30 onto the old SUM path.
|
||||
@@ -0,0 +1,357 @@
|
||||
# M6 — PHP-FPM, Memory, Swap & Worker Capacity
|
||||
|
||||
```text
|
||||
STATUS: COMPLETE (read-only; no production changes)
|
||||
WHEN: 2026-08-23 15:29–15:32 Europe/Ljubljana (UTC 13:29)
|
||||
HOST: server3, Debian 13, kernel 6.12.96, 8 cores, uptime 24d 18h
|
||||
```
|
||||
|
||||
Nothing was restarted, tuned, flushed, or edited on server3.
|
||||
|
||||
---
|
||||
|
||||
## Verdict
|
||||
|
||||
```text
|
||||
Swap: KEEP 2 GiB. Nearly full, but NOT thrashing (si/so=0, PSI memory=0).
|
||||
PHP-FPM max_children=14: KEEP. Peak has used all 14; listen queue never built.
|
||||
Do not raise FPM/Horizon on this host until RAM is freed or the machine is larger.
|
||||
InnoDB 10 GB pool: KEEP for 30-day snapshot growth (~7.5 GB table). Currently ~24% filled.
|
||||
OPcache 256 MB: KEEP (runtime used unknown; config is not tight on paper).
|
||||
SSR: ~116 MB RSS; not a memory problem. Frequent SIGTERM restarts are operator/supervisor, not OOM.
|
||||
Redis: 879 MB used / 2 GB max; RSS 167 MB — likely swapped. 0 evictions.
|
||||
```
|
||||
|
||||
Largest memory consumers are **MySQL (8.1 GB RSS)**, **vision Docker/Python (~2.5 GB)**, **ClamAV (1.1 GB)** — not Skinbase FPM.
|
||||
|
||||
---
|
||||
|
||||
## 1. Server RAM / swap
|
||||
|
||||
| Metric | Value |
|
||||
| ------ | ----- |
|
||||
| MemTotal | 23 GiB (24,615,736 kB) |
|
||||
| Mem used | 15 GiB |
|
||||
| MemAvailable | **8.2 GiB** |
|
||||
| Buffers+Cached | 7.2 GiB |
|
||||
| AnonPages | 13.4 GiB |
|
||||
| SwapTotal | 2.0 GiB |
|
||||
| Swap used | **2.0 GiB (37 MiB free)** |
|
||||
| SwapCached | 774 MiB |
|
||||
| Committed_AS | 30.1 GiB |
|
||||
| vm.swappiness | 10 |
|
||||
| vm.overcommit_memory | 0 |
|
||||
| Load | 5.09 / 7.07 / 8.33 (8 CPUs) |
|
||||
| PSI memory | some/full avg10=**0.00** |
|
||||
| PSI cpu | some avg10=4.7, avg60=14.5 |
|
||||
| vmstat si/so (1s×5) | **0 / 0** |
|
||||
| Processes | 243 |
|
||||
|
||||
**Swap interpretation:** cold / previously used pages sitting in swap, with **774 MiB SwapCached**. Not active thrashing. Do **not** disable swap. Do **not** resize without a RAM plan. swappiness=10 is already conservative.
|
||||
|
||||
---
|
||||
|
||||
## 2. Process memory (RSS, measured)
|
||||
|
||||
| Process | RSS | Notes |
|
||||
| ------- | --: | ----- |
|
||||
| mysqld | **8087 MB** | InnoDB pool configured 10 GB |
|
||||
| uvicorn/python (vision, port 8000) | **1310 + 542 + 238 + 197 + 103 + 54 MB** | Docker vision stack |
|
||||
| clamd | **1107 MB** | |
|
||||
| qdrant | 262 MB | vision |
|
||||
| meilisearch | 395 MB | Skinbase search |
|
||||
| crowdsec | 296 MB | |
|
||||
| redis-server | 171 MB RSS / **879 MB** `used_memory` | swapped |
|
||||
| netdata | 180 MB | |
|
||||
| gitea | 156 MB | |
|
||||
| php-fpm master | 48 MB | |
|
||||
| php-fpm pool skinbase (×6) | **88–110 MB** each | see §4 |
|
||||
| php-fpm pool www (×2) | 13 MB each | unused by Skinbase vhost |
|
||||
| horizon master | 104 MB | |
|
||||
| horizon supervisors (×3) | ~103 MB each | |
|
||||
| horizon workers (×6 now) | 103–125 MB | `--memory=128` |
|
||||
| node SSR | 116 MB | uptime minutes (restarted) |
|
||||
| reverb | 106 MB | |
|
||||
| nginx workers | 10–58 MB | one “shutting down” 15d |
|
||||
| php8.4-fpm cgroup | current **746 MB**, peak **1575 MB** | all pools |
|
||||
|
||||
---
|
||||
|
||||
## 3. PHP-FPM Skinbase pool (`php8.4-fpm-skinbase.sock`)
|
||||
|
||||
From `/etc/php/8.4/fpm/pool.d/skinbase.conf`:
|
||||
|
||||
| Setting | Production value |
|
||||
| ------- | ---------------- |
|
||||
| pm | **dynamic** |
|
||||
| pm.max_children | **14** |
|
||||
| pm.start_servers | 4 |
|
||||
| pm.min_spare_servers | 2 |
|
||||
| pm.max_spare_servers | 6 |
|
||||
| pm.process_idle_timeout | unset (N/A for dynamic) |
|
||||
| pm.max_requests | **500** |
|
||||
| request_terminate_timeout | 120s |
|
||||
| request_slowlog_timeout | 10s |
|
||||
| slowlog | `/var/log/php8.4-fpm-skinbase-slow.log` |
|
||||
| pm.status_path | `/fpm-status` (localhost) |
|
||||
| ping.path | `/fpm-ping` |
|
||||
| listen.backlog | unset → PHP default **511** |
|
||||
| memory_limit | **256M** (pool admin) |
|
||||
| max_execution_time | 60 |
|
||||
|
||||
Separate `www` pool: `pm.max_children=5` on `/run/php/php8.4-fpm.sock`.
|
||||
|
||||
---
|
||||
|
||||
## 4. FPM status (localhost `/fpm-status`, pool up ~28.7h since 2026-08-22 10:50)
|
||||
|
||||
| Field | Value |
|
||||
| ----- | ----- |
|
||||
| accepted conn | 214,108 |
|
||||
| listen queue | **0** |
|
||||
| max listen queue | **0** |
|
||||
| listen queue len | 0 |
|
||||
| idle processes | 5 |
|
||||
| active processes | 1 |
|
||||
| total processes | 6 |
|
||||
| max active processes | **14** |
|
||||
| max children reached | **10** |
|
||||
| slow requests (since start) | **2** |
|
||||
| memory peak (FPM counter) | 60 MB |
|
||||
|
||||
**Not currently exhausting workers.** Peak has used all 14 children. No socket backlog was recorded, so nginx did not sit behind a full listen queue — capacity is **tight at peak, idle most of the time**.
|
||||
|
||||
Worker RSS (KB): 87712, 89996, 106444, 109308, 109476, 109480.
|
||||
|
||||
| Stat | RSS |
|
||||
| ---- | --: |
|
||||
| min | 86 MB |
|
||||
| median | **108 MB** |
|
||||
| mean | **102 MB** |
|
||||
| max | **110 MB** |
|
||||
| PSS | not readable (permission) |
|
||||
| oldest worker | ~7.5 min (max_requests=500 recycling) |
|
||||
|
||||
`memory_limit` 256M is the ceiling, not typical RSS. **Do not size max_children from 256M × 14.**
|
||||
|
||||
Safe children from measured RSS:
|
||||
|
||||
```text
|
||||
typical_rss ≈ 110 MB
|
||||
14 × 110 MB ≈ 1.54 GB (matches systemd MemoryPeak 1.58 GB for all FPM)
|
||||
14 × 256 MB ≈ 3.58 GB worst-case if every worker hits the PHP limit
|
||||
```
|
||||
|
||||
Raising max_children on a host with **swap already full** and AnonPages 13.4 GiB is not justified. **Keep 14.**
|
||||
|
||||
---
|
||||
|
||||
## 5. OPcache (config; FPM runtime stats not exposed)
|
||||
|
||||
| Setting | Value |
|
||||
| ------- | ----- |
|
||||
| enable | 1 |
|
||||
| memory_consumption | 256 MB |
|
||||
| interned_strings_buffer | 32 MB |
|
||||
| max_accelerated_files | 50,000 |
|
||||
| validate_timestamps | 1 |
|
||||
| revalidate_freq | 2 s |
|
||||
| JIT | disable / 0 |
|
||||
|
||||
CLI cannot read the FPM cache. 50k file slots and 256 MB are large for this Laravel app. **Do not increase.** Optional later (MEDIUM, not memory): `validate_timestamps=0` after atomic releases.
|
||||
|
||||
---
|
||||
|
||||
## 6. Horizon (Supervisor `skinbase-horizon`)
|
||||
|
||||
Production env (`config/horizon.php` on the server):
|
||||
|
||||
| Supervisor | queues | maxProcesses | timeout | tries | memory |
|
||||
| ---------- | ------ | ------------: | ------: | ----: | -----: |
|
||||
| supervisor-default | search, default | 5 | 960 | 1 | 128 |
|
||||
| supervisor-messaging | broadcasts, notifications | 3 | 90 | 1 | 128 |
|
||||
| supervisor-mail | mail | 2 | 90 | 5 | 128 |
|
||||
|
||||
Measured now: 1 master + 3 supervisors + 6 workers ≈ **1.15 GB RSS**.
|
||||
|
||||
Peak if all maxProcesses spawn: 1+3+5+3+2 = **14 PHP procs × ~110 MB ≈ 1.5 GB**.
|
||||
|
||||
Mail isolation is intact (`tries=5`, own supervisor). **Do not merge queues.**
|
||||
|
||||
Redis lists (prefix as used by Laravel):
|
||||
|
||||
| List | LLEN |
|
||||
| ---- | ---: |
|
||||
| queues:default | **796** |
|
||||
| queues:mail | 0 |
|
||||
| queues:search | 0 |
|
||||
| queues:broadcasts | 0 |
|
||||
| queues:notifications | 0 |
|
||||
|
||||
**HIGH:** default queue depth 796 — worker **throughput**, not FPM RAM. Investigate job mix in a later milestone; do not add Horizon processes until RAM is free (each extra worker ≈ 110 MB and `--memory=128` is already near RSS 125 MB on default).
|
||||
|
||||
Horizon last supervisor restart: 2026-08-23 14:36 (clean exit 0 / SIGTERM wait), not OOM.
|
||||
|
||||
---
|
||||
|
||||
## 7. SSR
|
||||
|
||||
- RSS **116 MB**, Node `/opt/www/virtual/SkinbaseNova/bootstrap/ssr/ssr.js`
|
||||
- Supervisor restarts today: 14:22, 14:30, 14:47, 15:14, 15:16, 15:28 — all **SIGTERM** “waiting to stop”, then spawn. Not crash loops from memory.
|
||||
- Heap not sampled (would require attaching to Node).
|
||||
- **Not material** to host memory pressure.
|
||||
|
||||
---
|
||||
|
||||
## 8. Redis (read-only via Laravel)
|
||||
|
||||
| Field | Value |
|
||||
| ----- | ----- |
|
||||
| used_memory | 879 MB (peak 917 MB) |
|
||||
| used_memory_rss | **167 MB** |
|
||||
| maxmemory | 2.00 GB |
|
||||
| policy | allkeys-lru |
|
||||
| fragmentation_ratio | **0.19** |
|
||||
| evicted_keys | **0** |
|
||||
| expired_keys | 225,783 |
|
||||
| connected_clients | 28 |
|
||||
| blocked_clients | 0 |
|
||||
| keys | 3773 (3026 with TTL) |
|
||||
| AOF | off |
|
||||
| ops/sec | ~50 |
|
||||
| hit/miss | 1.46M / 1.39M |
|
||||
|
||||
frag 0.19 + RSS << used_memory ⇒ **Redis pages are in swap**. Latency risk under cache bursts. 0 evictions: cache is within 2 GB. Do not flush. Do not raise maxmemory.
|
||||
|
||||
---
|
||||
|
||||
## 9. MySQL / Percona (read-only)
|
||||
|
||||
| Field | Value |
|
||||
| ----- | ----- |
|
||||
| innodb_buffer_pool_size | **10.00 GB** (10 instances) |
|
||||
| pages_data / total | 156,554 / 655,360 = **23.9%** |
|
||||
| pages_free | 498,726 |
|
||||
| pool hit | 1 − 104,104 / 23,840,832,084 ≈ **99.9996%** |
|
||||
| max_connections | 120 |
|
||||
| Max_used_connections | 26 |
|
||||
| Threads_connected / running | 12 / 3 |
|
||||
| tmp tables / disk tmp | 1,987,237 / **21** |
|
||||
| innodb_redo_log_capacity | 2 GB |
|
||||
| Slow_queries | 11,997 (uptime 103,663 s) |
|
||||
|
||||
10 GB pool is **oversized for today’s ~2.8 GB schema**, but M5 30-day hourly snapshots grow toward **~7.5 GB**. Keeping 10 GB avoids refitting later. The table does **not** need to sit entirely in the pool; hit ratio is already excellent. **Do not shrink now; do not grow.**
|
||||
|
||||
---
|
||||
|
||||
## 10. Combined memory budget (measured / plausible peak)
|
||||
|
||||
| Bucket | Steady now | Plausible peak |
|
||||
| ------ | ---------: | -------------: |
|
||||
| OS/page cache (reclaimable) | ~7 GB cache | shrinks under pressure |
|
||||
| MySQL | 8.1 GB | ~10 GB pool |
|
||||
| Vision Docker/Python/Qdrant | ~2.7 GB | similar |
|
||||
| ClamAV | 1.1 GB | similar |
|
||||
| Meilisearch | 0.40 GB | 0.5 GB |
|
||||
| Crowdsec | 0.30 GB | similar |
|
||||
| Redis RSS | 0.17 GB | 2 GB if unswapped |
|
||||
| PHP-FPM Skinbase | 0.65 GB (6×110) | **1.54 GB** (14×110) |
|
||||
| PHP-FPM www | 0.03 GB | 0.06 GB |
|
||||
| Horizon | 1.15 GB | **1.5 GB** |
|
||||
| SSR + Reverb | 0.22 GB | 0.3 GB |
|
||||
| nginx | 0.20 GB | 0.25 GB |
|
||||
| Netdata/Gitea/journald/docker | ~0.5 GB | similar |
|
||||
| Swap | 2 GB **in use** | — |
|
||||
|
||||
Steady anonymous ~14 GB + 2 GB swap explains the full swap file while **8.2 GB still Available** as cache. Peak Skinbase (FPM 14 + Horizon 14) ≈ **3.0 GB**, already observed in FPM cgroup peak 1.58 GB.
|
||||
|
||||
---
|
||||
|
||||
## 11. PHP-FPM recommendation
|
||||
|
||||
**NO CHANGE** to pool sizing.
|
||||
|
||||
```text
|
||||
pm = dynamic
|
||||
pm.max_children = 14 # keep; peak already 14, listen queue 0
|
||||
pm.start_servers = 4 # keep
|
||||
pm.min_spare_servers = 2 # keep
|
||||
pm.max_spare_servers = 6 # keep (matches current total=6)
|
||||
pm.max_requests = 500 # keep; workers recycle in minutes
|
||||
```
|
||||
|
||||
Calculation: median RSS 108 MB; 14×108≈1.5 GB. Available RAM is cache, not idle anonymous. Raising children increases swap of Redis/MySQL. Lowering below 14 would clip the measured peak (`max active processes=14`).
|
||||
|
||||
---
|
||||
|
||||
## 12. Swap recommendation
|
||||
|
||||
| Option | Decision |
|
||||
| ------ | -------- |
|
||||
| Keep 2 GiB swap | **YES** |
|
||||
| Resize swap | **NO** (overflow is working; host needs RAM or fewer co-tenants) |
|
||||
| Change vm.swappiness | **NO** (already 10) |
|
||||
| Disable swap | **NO** |
|
||||
|
||||
---
|
||||
|
||||
## 13. OOM / crash safety
|
||||
|
||||
| Check | Result |
|
||||
| ----- | ------ |
|
||||
| Kernel OOM (`journalctl -k` / dmesg) | **no lines** visible to this user |
|
||||
| PHP-FPM worker memory | **YES** — `Allowed memory size of 268435456` Predis `StreamConnection.php` 2026-08-13 and 2026-08-18 |
|
||||
| Redis OOM / evictions | **no** (evicted_keys=0) |
|
||||
| MySQL | no evidence in readable logs |
|
||||
| Horizon | clean supervisor stop 14:36, not killed |
|
||||
| SSR | SIGTERM restarts, not OOM |
|
||||
|
||||
---
|
||||
|
||||
## 14. Ranked findings
|
||||
|
||||
| Sev | Finding | Change now? |
|
||||
| --- | ------- | ----------- |
|
||||
| HIGH | Swap 2 GiB ~full; Redis RSS≪used_memory | No sysctl. Free RAM (vision/clam) or larger VM later |
|
||||
| HIGH | `queues:default` LLEN=**796** | Not FPM. Later queue milestone; don’t add Horizon RAM |
|
||||
| HIGH | Shared host: vision ~2.7 GB + clamd 1.1 GB | Out of Skinbase app; ops placement |
|
||||
| MEDIUM | FPM `max_children` reached 10×, max active 14 | Keep 14; watch listen queue |
|
||||
| MEDIUM | PHP 256 MB fatals on Predis | App/job payload size, not pool size |
|
||||
| MEDIUM | slowlog 10s, 139k historical lines; status slow=2 since 22 Aug | Leave timeout; log rotation ops |
|
||||
| LOW | nginx worker “shutting down” 15d | Recycle nginx in a maintenance window |
|
||||
| LOW | www pool unused by Skinbase vhost | Optional disable later |
|
||||
| NO CHANGE | OPcache 256/32/50k | |
|
||||
| NO CHANGE | InnoDB 10 GB | Needed as snapshots grow |
|
||||
| NO CHANGE | Horizon isolation / mail supervisor | |
|
||||
| NO CHANGE | SSR RSS | |
|
||||
|
||||
---
|
||||
|
||||
## 15. Expected benefit of doing nothing to FPM
|
||||
|
||||
Avoid swapping MySQL/Redis further. Current request path is not FPM-bound (`active=1`, `listen queue=0`).
|
||||
|
||||
---
|
||||
|
||||
## 16. Deployment plan (only if a change is later approved)
|
||||
|
||||
M6 recommends **no FPM/swap/sysctl/MySQL/Horizon count change**.
|
||||
|
||||
If ops later moves vision off-box or adds RAM:
|
||||
|
||||
1. Re-measure RSS and `/fpm-status`.
|
||||
2. Then consider `pm.max_children` using **new** median RSS, not 256M.
|
||||
3. Pool edit + `systemctl reload php8.4-fpm` (reload, not restart if possible).
|
||||
4. Rollback: restore `skinbase.conf` and reload.
|
||||
|
||||
Any `pm.*` or `memory_limit` change **requires FPM reload**. Swap/sysctl/MySQL buffer pool need service-level restarts — not proposed.
|
||||
|
||||
---
|
||||
|
||||
## 17. Risks of increasing max_children without more RAM
|
||||
|
||||
More anonymous PHP, more swap of Redis (cache latency) and InnoDB (buffer pool eviction), possible FPM 256 MB fatals still happen per-request.
|
||||
|
||||
---
|
||||
|
||||
Production was not modified.
|
||||
@@ -0,0 +1,186 @@
|
||||
# M7 — Default Queue Backlog & Redis/PHP Memory Failures
|
||||
|
||||
```text
|
||||
STATUS: COMPLETE (investigation + local smallest fixes; production unchanged)
|
||||
WHEN: 2026-08-23 17:49 Europe/Ljubljana
|
||||
```
|
||||
|
||||
No Horizon restart, Redis flush, job retry, memory_limit change, or worker-count change on server3.
|
||||
|
||||
---
|
||||
|
||||
## Verdict
|
||||
|
||||
```text
|
||||
Default backlog class: C + D (producer > consumer; ranking/rec jobs block index jobs)
|
||||
Default depth: 1208 → 1213 in 60s (net +5/min, growing)
|
||||
Oldest default job: ~2.5 h
|
||||
Search queue: 0 (idle workers exist on supervisor-default)
|
||||
collections: 32,632 jobs, oldest 2026-04-30, NO Horizon consumer
|
||||
forum-moderation: 552,888 AnalyzeForumPostJob, ~557 MB list, NO consumer
|
||||
forum-security: 5,533 jobs, no consumer
|
||||
presence index: SET 2,480,292 members, ~272 MB, TTL -1
|
||||
Predis 256 MB fatals: loading those huge Redis replies over HTTP/FPM
|
||||
Horizon workers: do NOT increase
|
||||
search+default supervisor: KEEP together; put Index* jobs on search
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 1. `queues:default` contents (metadata only)
|
||||
|
||||
| Class | Count | Avg payload |
|
||||
| ----- | ----: | ----------: |
|
||||
| `RankBuildScopeListsJob` | **588** | 1004 B |
|
||||
| `IndexUserJob` | **395** | 954 B |
|
||||
| `RecComputeSimilarByTagsJob` | **169** | 1212 B |
|
||||
| `IndexArtworkJob` | 40 | 918 B |
|
||||
| `RecalculateRisingNovaCardsJob` | 10 | 987 B |
|
||||
| `RebuildTrendingNovaCardsJob` | 3 | 977 B |
|
||||
| `RankBuildListsJob` | 2 | 931 B |
|
||||
| `RecBuildItemPairsFromFavouritesJob` | 1 | 1102 B |
|
||||
|
||||
- Oldest pushedAt: **2026-08-23 15:19** (~2.5 h)
|
||||
- Newest: 17:48 (~75 s)
|
||||
- Delayed: 1 `RankBuildScopeListsJob`
|
||||
- Reserved: 0
|
||||
- Attempts: 1206×0, 2×1
|
||||
- Max payload: **1212 B** — not oversized job payloads
|
||||
|
||||
Net 60s: 1208 → 1213 (**accumulating**).
|
||||
|
||||
---
|
||||
|
||||
## 2. Throughput
|
||||
|
||||
Horizon metrics repo empty from this user (prefix split `skinbase_horizon` vs `skinbasenova_horizon`).
|
||||
|
||||
Evidence: backlog growing +5 net/min while 5 default-supervisor processes exist. Ranking jobs `timeout=300s`; rec tag jobs `timeout=600s`. A few long jobs starve `IndexUserJob` (30s, user-facing search).
|
||||
|
||||
---
|
||||
|
||||
## 3. Producers
|
||||
|
||||
| Producer | Job | Queue | Cadence | Volume |
|
||||
| -------- | --- | ----- | ------- | ------ |
|
||||
| `RankBuildListsJob` hourlyAt(15) | `RankBuildScopeListsJob` per category/content type | default | hourly fan-out | **~588/hour** still sitting |
|
||||
| `UserStatsService::reindex` | `IndexUserJob` | **default** (bug; Scout config is `search`) | per stats change | 395 queued |
|
||||
| Artwork indexers | `IndexArtworkJob` | **default** (same bug) | publish/trending/tags | 40 queued |
|
||||
| Scheduler rec nightly | `RecComputeSimilarByTagsJob` batches | default (`RECOMMENDATIONS_QUEUE`) | nightly + leftover | 169 |
|
||||
| Nova card crons | Rising/Trending rebuild | default | 15 min / hourly | 13 |
|
||||
| `collections:dispatch-maintenance` hourlyAt(43) | Health/rec/dup | **collections** | hourly ~110 | **32k unconsumed** |
|
||||
| Forum plugin | `AnalyzeForumPostJob` | **forum-moderation** | continuous | **552,888 unconsumed** |
|
||||
|
||||
---
|
||||
|
||||
## 4. Isolation
|
||||
|
||||
Move **IndexArtworkJob / IndexUserJob** to `search` (already on supervisor-default, currently idle). Same worker budget, better latency. Do **not** add processes.
|
||||
|
||||
Do **not** merge mail. Rec stays on default until ranking unique + M1 afterArtworkId is healthy.
|
||||
|
||||
**collections** and **forum-moderation** have **no Horizon supervisor** — isolation without a consumer is a dump. Do not add workers until RAM is freed (M6). Stop/limit dispatch first.
|
||||
|
||||
---
|
||||
|
||||
## 5. Predis 256 MB fatals (13 Aug, 18 Aug)
|
||||
|
||||
FPM error log (Skinbase pool, HTTP):
|
||||
|
||||
```text
|
||||
Allowed memory size of 268435456 bytes exhausted
|
||||
(tried to allocate 67108872 bytes)
|
||||
vendor/predis/predis/src/Connection/StreamConnection.php:212
|
||||
```
|
||||
|
||||
Two Redis values large enough to explode a 256 MB FPM worker when Predis reads the reply:
|
||||
|
||||
| Key | Type | Size | How it is read |
|
||||
| --- | ---- | ---- | -------------- |
|
||||
| `queues:forum-moderation` | list | **552,888 jobs ≈ 557 MB** | Horizon UI / any `LRANGE` of the queue |
|
||||
| `skinbase:presence:online:index` | set | **2,480,292 members ≈ 272 MB** | `OnlineVisitorRepository::readIndexMembers()` → **`SMEMBERS`** then `all()` on moderation traffic page |
|
||||
|
||||
Index members have **TTL -1**; per-visitor keys TTL 300s. Expired visitors are never removed unless `all()` finishes — it cannot, so the set grows forever.
|
||||
|
||||
**Do not raise memory_limit.** Stop loading the full set/list in one Redis command.
|
||||
|
||||
Legacy prefix `skinbasenova-database-queues:*` still holds leftover lists (another ~11 MB default + forum copies).
|
||||
|
||||
---
|
||||
|
||||
## 6. Payload size
|
||||
|
||||
Default-queue jobs are **~1 KB**. Not the memory issue.
|
||||
|
||||
---
|
||||
|
||||
## 7. Largest Redis keys/prefixes (SCAN, names redacted)
|
||||
|
||||
| Group | Approx |
|
||||
| ----- | -----: |
|
||||
| `queues:forum-moderation` | 557 MB |
|
||||
| `skinbase:presence:online:index` | 272 MB |
|
||||
| `queues:collections` | 35 MB |
|
||||
| `skinbasenova-database-queues:default` (legacy prefix) | 11 MB |
|
||||
| `queues:forum-security` | 5.5 MB |
|
||||
| Horizon recent/pending (legacy prefix) | ~3 MB |
|
||||
|
||||
Keyspace ~3.7k keys but a few lists/sets dominate `used_memory` (~879 MB).
|
||||
|
||||
---
|
||||
|
||||
## 8. search vs default supervisor
|
||||
|
||||
**Keep sharing supervisor-default.** Search workers are idle because index jobs were on `default`. After Index* → `search`, auto-balance can drain indexing without extra RAM.
|
||||
|
||||
---
|
||||
|
||||
## 9. Backlog classification
|
||||
|
||||
**Default: C+D** — hourly ranking fan-out + 10-minute rec jobs exceed 5 workers; index jobs wait behind them. Not a retry storm (attempts mostly 0).
|
||||
|
||||
**collections / forum-*: F** — orphan queues, growing for months (collections from 2026-04-30; forum still enqueuing today).
|
||||
|
||||
**failed_jobs:** 133 rows, **all today 14:37–14:56** (Horizon SIGTERM 14:36). Classes: RecCompute* `Typed property $afterArtworkId must not be accessed before initialization` on release `20260823-12*` — production rec job constructor mismatch. Enhance jobs: curl 127.0.0.1:8095 refused. **Do not retry.**
|
||||
|
||||
---
|
||||
|
||||
## 10. User-facing latency
|
||||
|
||||
`IndexUserJob` / `IndexArtworkJob` wait behind ranking/rec on default. Oldest index jobs share the 2.5 h window. Emails/notifications not on default (mail/notify queues empty). Uploads/vision queue empty.
|
||||
|
||||
---
|
||||
|
||||
## 11. Redis latency vs swap
|
||||
|
||||
ops/sec ~42, blocked clients 0, evicted 0. Slowlog parse inconclusive. Swapped Redis **will** hurt when someone reads a 557 MB list (that *is* the 256 MB fatal). Queue pop of 1 KB jobs is small.
|
||||
|
||||
---
|
||||
|
||||
## 12. Smallest fix (local; not deployed)
|
||||
|
||||
1. `IndexArtworkJob` + `IndexUserJob` → `scout.queue.queue` (`search`) + `ShouldBeUnique`.
|
||||
2. `RankBuildScopeListsJob` `ShouldBeUnique` per scope so hourly fan-out cannot stack 588 copies.
|
||||
3. Presence index: **SSCAN + cap 2000**, never `SMEMBERS` of 2.4M.
|
||||
|
||||
**Not in this deploy:** delete 552k forum jobs, add Horizon supervisors, raise PHP memory, raise Horizon processes.
|
||||
|
||||
**After RAM (M6):** optional `supervisor-collections` / forum workers, or stop those dispatchers.
|
||||
|
||||
---
|
||||
|
||||
## 13. Tests
|
||||
|
||||
`tests/Feature/Queue/DefaultQueueAssignmentTest.php`
|
||||
|
||||
---
|
||||
|
||||
## 14. Deploy
|
||||
|
||||
Ship the three job/presence changes. Reload Horizon so new job `onQueue` applies to **new** dispatches (existing 395 IndexUserJob stay on default until drained). No FPM memory_limit change. No Redis delete.
|
||||
|
||||
Rollback: revert those files; Horizon reload.
|
||||
|
||||
Risk: unique ranking skips a scope if a stale unique lock remains (`uniqueFor=360s`). Presence admin UI shows at most 2000 index members until a later prune job (write) is designed.
|
||||
|
||||
Production not modified.
|
||||
@@ -0,0 +1,217 @@
|
||||
# M7.1 — Queue Compatibility & Backlog Containment
|
||||
|
||||
```text
|
||||
STATUS: COMPLETE + FINALIZED (local; production unchanged)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 1. RecCompute serialization root cause
|
||||
|
||||
M1 added constructor-promoted `private readonly ?int $afterArtworkId = null`.
|
||||
|
||||
**Constructor defaults are not applied on `unserialize`.** PHP 8 leaves a **typed property uninitialized** if it is missing from the serialized payload. First access (handle, failed(), logging) throws:
|
||||
|
||||
```text
|
||||
Typed property App\Jobs\RecComputeSimilar*::$afterArtworkId must not be accessed before initialization
|
||||
```
|
||||
|
||||
Old queued jobs (pre-M1, two constructor args) and Laravel `SerializesModels` (skips default-null properties on serialize) both omit the key.
|
||||
|
||||
**Fix:** class-level `private ?int $afterArtworkId = null` (not readonly promotion) + `__unserialize` calls `isset` restore. Getter `afterArtworkId()` never reads uninitialized state.
|
||||
|
||||
Cursor batching unchanged: `WHERE id > afterArtworkId`, `ORDER BY id`, `LIMIT batch`, no OFFSET.
|
||||
|
||||
---
|
||||
|
||||
## 2. Stale Tags jobs — worst case (live 2026-08-23 18:xx)
|
||||
|
||||
Live `queues:default` RecCompute* (not failed_jobs):
|
||||
|
||||
- **Behavior: 0 waiting**
|
||||
- **Hybrid: 0 waiting**
|
||||
- **Tags: 169 waiting**, all `artworkId=null` (batch mode), **all already have `afterArtworkId` set** (cursors 1284–4517). These are overlapping remainder chains, not missing-property payloads.
|
||||
|
||||
What `afterArtworkId=null` would do: process first 200 public artworks ordered by id, then `dispatch(null, 200, lastId)` until the table is exhausted. One full rebuild ≈ `ceil(N/200)` jobs. With ~50k snapshot-eligible artworks that is **~250 jobs per chain**.
|
||||
|
||||
If all 169 were null-cursor roots: **169 × ~250 ≈ 42,000 tag jobs** (duplicate full rebuilds). That is the theoretical worst case.
|
||||
|
||||
**Actual live worst case:** they continue from mid-table. Example: 47 jobs with `after=3709` each walk the remainder (~232 batches) → **~11,000 overlapping jobs from that cursor group alone**, plus the other cursor groups. Still dangerous.
|
||||
|
||||
**Do not clear `queues:default` wholesale.** Operator command (not scheduled):
|
||||
|
||||
```bash
|
||||
php artisan skinbase:purge-queued-rec-tags --queue=default --dry-run
|
||||
php artisan skinbase:purge-queued-rec-tags --queue=default --execute
|
||||
```
|
||||
|
||||
Removes **waiting** `RecComputeSimilarByTagsJob` only. Preserves ranking/index/other. Does not touch reserved or `failed_jobs`. Then tonight’s scheduler dispatches **one** clean Tags rebuild (`afterArtworkId=null` once).
|
||||
|
||||
---
|
||||
|
||||
## 3. Failed-job breakdown (complete, n=133, all queue=default)
|
||||
|
||||
M7 sampled 80 rows. Full table:
|
||||
|
||||
| Class | Count | Cause |
|
||||
| ----- | ----: | ----- |
|
||||
| RecComputeSimilarHybridJob | **66** | `$afterArtworkId` uninitialized |
|
||||
| RecComputeSimilarByBehaviorJob | **42** | same |
|
||||
| RecComputeSimilarByTagsJob | **10** | same |
|
||||
| Enhance\ProcessEnhanceJob | **8** | curl 127.0.0.1:8095 refused |
|
||||
| App\Mail\AcademyAccessIssue | **7** | mail on default |
|
||||
| **Total** | **133** | |
|
||||
|
||||
66+42+10+8+7=133. Do not retry.
|
||||
|
||||
---
|
||||
|
||||
## 4. Index* queue
|
||||
|
||||
`IndexArtworkJob` / `IndexUserJob` use `config('scout.queue.queue')` (default `search`).
|
||||
|
||||
`ShouldBeUniqueUntilProcessing`:
|
||||
|
||||
| | IndexUser | IndexArtwork |
|
||||
| --- | --- | --- |
|
||||
| uniqueId | userId | artworkId |
|
||||
| uniqueFor | 60s | 120s |
|
||||
|
||||
handle() loads the model from DB, so a dropped duplicate while queued still indexes **current** stats. After the worker **starts**, a later change can enqueue again (not lost). Dead-worker lock expires at uniqueFor.
|
||||
|
||||
---
|
||||
|
||||
## 5. Ranking dedupe
|
||||
|
||||
`RankBuildScopeListsJob` `ShouldBeUnique`, `uniqueId = scopeType:scopeId`.
|
||||
|
||||
**uniqueFor=21600 (6 hours)** from `ranking.scope_job_unique_for` / `RANK_SCOPE_JOB_UNIQUE_FOR`. 360s expired long before a 2.5h wait, so hourly fan-out stacked duplicates. 6h > current wait + 300s timeout; dead locks recover the same day. Different scopes stay independent.
|
||||
|
||||
---
|
||||
|
||||
## 6. Presence memory
|
||||
|
||||
`SMEMBERS` removed. `SSCAN` + cap `traffic.online_visitors.index_read_limit` (default 2000). Admin UI may under-count vs 2.48M stale members until prune.
|
||||
|
||||
### Stale-member prune (design only — not executed)
|
||||
|
||||
1. Detect: SSCAN members; `GET skinbase:presence:online:{member}`; if missing/expired JSON, member is stale.
|
||||
2. Batches: SSCAN COUNT 200; `SREM` up to 200 stale ids; sleep 50ms.
|
||||
3. Cost: 2.48M SSCAN + GET. At 200/batch ~12k round trips. Off-peak 5–15 min. Do not `DEL` the index key (would drop live members).
|
||||
4. Cadence: hourly `withoutOverlapping`, stop after N batches/run.
|
||||
|
||||
---
|
||||
|
||||
## 7. Orphan producers
|
||||
|
||||
| Queue | Consumer? | Enabled? | Useful if months late? | Dispatch today? | M7.1 |
|
||||
| ----- | --------- | -------- | --------------------- | --------------- | ---- |
|
||||
| collections | **No** Horizon supervisor | scheduler hourlyAt(43) | Refresh is idempotent; delayed run OK | yes until gate | **`COLLECTIONS_V5_DISPATCH_ENABLED=false`** (default) |
|
||||
| forum-moderation | **No** Horizon supervisor | gated | **No** for months-old posts | was yes | **`dispatchAsyncScan()` now returns without queueing when `skinbase_ai_moderation.queue.dispatch_enabled` is false.** Topic create/edit and `forum:ai-scan` (async) are gated. `forum:scan-posts` and `forum:ai-scan --sync` stay in-process. |
|
||||
| forum-security | **No** | gated | stale in hours | 5.5k | **`forum:bot-scan` / `forum:firewall-scan` async dispatch gated.** `--sync` still runs. Scheduler omits those commands when dispatch is false. |
|
||||
|
||||
Do not add Horizon workers (M6 RAM).
|
||||
|
||||
---
|
||||
|
||||
## 8. Existing orphan jobs (do not delete)
|
||||
|
||||
| Queue | Class | Disposition |
|
||||
| ----- | ----- | ----------- |
|
||||
| collections | Health / recommendation / duplicate | **KEEP AND PROCESS LATER** or **SAFE TO DISCARD** (recompute from current rows). Product: either. |
|
||||
| forum-moderation | AnalyzeForumPostJob × 552,888 | **SAFE TO DISCARD** for age > 48h; **NEEDS PRODUCT DECISION** for last day. Analyzing May–August posts has no moderation value. |
|
||||
| forum-security | ~5.5k | **SAFE TO DISCARD** (burst/firewall windows expired). |
|
||||
|
||||
Later cleanup (not now): `UNLINK queues:forum-moderation` (async, non-blocking) after stopping producers; or `LTRIM`/`RPOP` in 5k batches. Never one giant `DEL` in a request.
|
||||
|
||||
---
|
||||
|
||||
## 9. Legacy prefix `skinbasenova-database-queues:*`
|
||||
|
||||
Current Laravel prefix is `skinbase-database-`. `skinbasenova-database-*` is leftover from an older `REDIS_PREFIX` / app name. Nothing in this codebase reads it. Horizon also has `skinbasenova_horizon:*` vs `skinbase_horizon:*`.
|
||||
|
||||
Reclaimable: ~11 MB default + ~9 MB forum + ~2.5 MB collections + Horizon recent lists (~3 MB) ≈ **25–30 MB** (not the 557 MB live forum list).
|
||||
|
||||
Cleanup only after confirming no process uses `skinbasenova-database-` (`HORIZON_PREFIX`, `REDIS_PREFIX` on all PHP/Horizon). Then `UNLINK` those keys. Do not delete `skinbase-database-queues:forum-moderation`.
|
||||
|
||||
---
|
||||
|
||||
## 10. Files changed
|
||||
|
||||
```text
|
||||
app/Jobs/Concerns/RestoresAfterArtworkIdCursor.php
|
||||
app/Jobs/RecComputeSimilarByTagsJob.php
|
||||
app/Jobs/RecComputeSimilarByBehaviorJob.php
|
||||
app/Jobs/RecComputeSimilarHybridJob.php
|
||||
app/Jobs/IndexArtworkJob.php
|
||||
app/Jobs/IndexUserJob.php
|
||||
app/Jobs/RankBuildScopeListsJob.php
|
||||
app/Services/Traffic/OnlineVisitorRepository.php
|
||||
app/Services/CollectionBackgroundJobService.php
|
||||
app/Console/Commands/DispatchCollectionMaintenanceCommand.php
|
||||
routes/console.php
|
||||
config/collections.php
|
||||
config/forum_security.php
|
||||
config/skinbase_ai_moderation.php
|
||||
config/traffic.php
|
||||
.env.example
|
||||
phpunit.xml
|
||||
tests/...
|
||||
docs/optimization-m7.1-queue-compatibility.md
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 11. Tests
|
||||
|
||||
15 passed: RecCompute unserialize, cursor batching, Index* search queue, ranking uniqueId per scope, presence SSCAN/no SMEMBERS, collection dispatch gate.
|
||||
|
||||
---
|
||||
|
||||
## 12. Deploy sequence (later; not now)
|
||||
|
||||
1. Ship this release (jobs, uniqueness, presence SSCAN, collection gate, purge command).
|
||||
2. Production `.env`:
|
||||
```text
|
||||
COLLECTIONS_V5_DISPATCH_ENABLED=false
|
||||
RANK_SCOPE_JOB_UNIQUE_FOR=21600
|
||||
```
|
||||
Do not rely on `FORUM_QUEUE_DISPATCH_ENABLED` until the Forum plugin is changed.
|
||||
3. `php artisan config:cache`
|
||||
4. `sudo supervisorctl restart skinbase-horizon`
|
||||
5. Optional: `sudo systemctl reload php8.4-fpm`
|
||||
6. **After Horizon is on the new code**, purge stale Tags (operator, not cron):
|
||||
```bash
|
||||
php artisan skinbase:purge-queued-rec-tags --queue=default --dry-run
|
||||
php artisan skinbase:purge-queued-rec-tags --queue=default --execute
|
||||
```
|
||||
7. Confirm `queues:default` RecComputeSimilarByTagsJob count is 0. Leave ranking/index jobs.
|
||||
8. Tonight’s scheduler may dispatch **one** Tags rebuild.
|
||||
|
||||
**Horizon restart: YES.** Do not retry `failed_jobs`. Do not `SMEMBERS` presence index. Do not `DEL` forum-moderation.
|
||||
|
||||
---
|
||||
|
||||
## 13. Post-deploy verification
|
||||
|
||||
```bash
|
||||
php artisan tinker --execute="echo json_encode([
|
||||
'default' => Redis::llen('queues:default'),
|
||||
'search' => Redis::llen('queues:search'),
|
||||
'collections' => Redis::llen('queues:collections'),
|
||||
'forum' => Redis::llen('queues:forum-moderation'),
|
||||
]);"
|
||||
# After some traffic, search LLEN should rise (Index* jobs), default ranking unique should stop stacking
|
||||
# Do not retry failed_jobs
|
||||
# Do not SMEMBERS skinbase:presence:online:index
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 14. Risks
|
||||
|
||||
- 169 in-queue tag jobs become “start from id 0” batches; duplicate compute, extra load, no crash.
|
||||
- Ranking unique skips a scope for up to 360s if a lock is stuck.
|
||||
- Presence UI shows ≤2000 members until prune.
|
||||
- Forum enqueue may continue if the plugin ignores `FORUM_QUEUE_DISPATCH_ENABLED`.
|
||||
- Collection tests set dispatch enabled in phpunit; production default is false.
|
||||
@@ -0,0 +1,180 @@
|
||||
# M8 — Redis Orphan Cleanup & Presence Compaction
|
||||
|
||||
```text
|
||||
STATUS: COMPLETE (tooling + tests; no production UNLINK/SREM executed)
|
||||
WHEN: 2026-08-23 ~18:55 Europe/Ljubljana
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 1. Production before snapshot (read-only)
|
||||
|
||||
| Metric | Value |
|
||||
| ------ | ----- |
|
||||
| used_memory | **883.41 MB** (926,320,256) |
|
||||
| used_memory_rss | **176.88 MB** |
|
||||
| used_memory_peak | 917.11 MB |
|
||||
| maxmemory | 2.00 GB |
|
||||
| fragmentation_ratio | **0.20** (RSS ≪ used → swapped) |
|
||||
| evicted_keys | 0 |
|
||||
| connected_clients | 28 |
|
||||
| blocked_clients | 0 |
|
||||
| ops/sec | 92 |
|
||||
| keys | 4838 (4100 with TTL) |
|
||||
| REDIS_PREFIX | `skinbase-database-` |
|
||||
| HORIZON_PREFIX | `skinbase_horizon:` |
|
||||
|
||||
| Queue | llen | delayed | reserved |
|
||||
| ----- | ---: | ------: | -------: |
|
||||
| default | 0 | 0 | 0 |
|
||||
| search | 0 | 0 | 0 |
|
||||
| mail | 0 | 0 | 0 |
|
||||
| collections | **32632** | 0 | 0 |
|
||||
| forum-moderation | **553090** | 0 | 0 |
|
||||
| forum-security | **5538** | 0 | 0 |
|
||||
|
||||
Presence `SCARD` **2,480,634**. Sample 200 members: **200 stale / 0 live** (ratio 1.0).
|
||||
|
||||
Legacy SCAN: `skinbasenova_horizon:` 662 keys ~7 MB; `skinbasenova-database-` 11 keys ~23 MB.
|
||||
|
||||
---
|
||||
|
||||
## 2. Producer containment
|
||||
|
||||
45s later: collections 32632, forum-moderation 553090, forum-security 5538 — **unchanged**.
|
||||
|
||||
Config: collections/forum_ai/forum_sec **dispatch_enabled=false**.
|
||||
|
||||
Newest forum payloads ~18:47 then **no growth**. **Proceed.**
|
||||
|
||||
---
|
||||
|
||||
## 3. Queue class samples (LINDEX/LRANGE ends only, not full 553k)
|
||||
|
||||
| Queue | Sample | Classes | Oldest | Newest | avg bytes |
|
||||
| ----- | -----: | ------- | ------ | ------ | --------: |
|
||||
| forum-moderation | 204 | **100% AnalyzeForumPostJob** | 2026-04-30 | 2026-08-23 18:47 | 987 |
|
||||
| forum-security | 204 | FirewallActivityMonitor + BotActivityMonitor | 2026-04-30 | 2026-08-23 18:47 | 993 |
|
||||
| collections | 204 | Duplicate / Health / Recommendation (~equal) | 2026-04-30 | 2026-08-23 17:43 | 1054 |
|
||||
|
||||
No reserved/delayed.
|
||||
|
||||
---
|
||||
|
||||
## 4. forum-moderation
|
||||
|
||||
**UNLINK waiting + notify** after dry-run.
|
||||
|
||||
Months-old `AnalyzeForumPostJob` has no moderation value. Command samples classes via LINDEX stride (no 553k LRANGE). Refuses if dispatch_enabled, reserved/delayed, or unexpected class (unless `--force`).
|
||||
|
||||
---
|
||||
|
||||
## 5. forum-security
|
||||
|
||||
**UNLINK.** Windowed monitors from April–August are expired. Same guards.
|
||||
|
||||
---
|
||||
|
||||
## 6. collections
|
||||
|
||||
**A. SAFE TO DISCARD AND REBUILD.** Jobs hold `collectionId` only; `handle()` loads current DB. Old queue is a 4-month backlog of the same refresh. Do not auto-delete; operator `--execute collections` when ready. Re-enable `COLLECTIONS_V5_DISPATCH_ENABLED` only after a Horizon consumer exists.
|
||||
|
||||
---
|
||||
|
||||
## 7. Presence stale ratio
|
||||
|
||||
Sample **100% stale** (200/200). Estimated ~2.48M ghost members, ~272 MB.
|
||||
|
||||
---
|
||||
|
||||
## 8–9. Presence pruning
|
||||
|
||||
`PresenceIndexPruner`: SSCAN, pipeline EXISTS, Lua `EXISTS record==0 then SREM`. Never SMEMBERS, never DEL index.
|
||||
|
||||
Race: Lua is atomic vs SETEX. If a visitor is rewritten after SREM, they SADD again on next track.
|
||||
|
||||
Schedule: **hourlyAt(33)** if `ONLINE_VISITOR_INDEX_PRUNE_ENABLED=true` (default **false**). Avoids :02 snapshot, :15 rank, :21 leaderboards, :25 nova, nightly rec 02:00. First catch-up: raise `--max-batches` manually (200 batches × 500 ≈ 100k members/run → ~25 hours at hourly default).
|
||||
|
||||
---
|
||||
|
||||
## 10–11. Legacy prefix
|
||||
|
||||
Current process uses `skinbase-database-` / `skinbase_horizon:`. Obsolete `skinbasenova-*` ~30 MB. Command uses an **unprefixed** Predis client, UNLINKs only `skinbasenova-database-` and `skinbasenova_horizon:` keys, never current prefixes.
|
||||
|
||||
---
|
||||
|
||||
## 12. Expected reclaim (used_memory, RSS later)
|
||||
|
||||
| Stage | ~used_memory |
|
||||
| ----- | -----------: |
|
||||
| forum-security | 5.5 MB |
|
||||
| forum-moderation | **557 MB** |
|
||||
| legacy prefix | 25–30 MB |
|
||||
| presence (after full prune) | **~272 MB** |
|
||||
| collections (if later) | 35 MB |
|
||||
| **Total if all stages** | **~870 MB** toward a small Redis |
|
||||
|
||||
RSS may stay ~177 MB until OS reclaims. No MEMORY PURGE.
|
||||
|
||||
---
|
||||
|
||||
## 13. Operator commands (not run here)
|
||||
|
||||
```bash
|
||||
# A audit
|
||||
php artisan tinker --execute="echo json_encode(['used'=>Redis::info('memory')['used_memory_human'],'mod'=>Redis::llen('queues:forum-moderation'),'sec'=>Redis::llen('queues:forum-security'),'col'=>Redis::llen('queues:collections'),'scard'=>Redis::scard('skinbase:presence:online:index')]);"
|
||||
|
||||
# B forum-security
|
||||
php artisan skinbase:redis-cleanup-orphans forum-security --dry-run
|
||||
php artisan skinbase:redis-cleanup-orphans forum-security --execute
|
||||
|
||||
# C forum-moderation
|
||||
php artisan skinbase:redis-cleanup-orphans forum-moderation --dry-run
|
||||
php artisan skinbase:redis-cleanup-orphans forum-moderation --execute
|
||||
|
||||
# D legacy
|
||||
php artisan skinbase:redis-cleanup-legacy-prefix --dry-run
|
||||
php artisan skinbase:redis-cleanup-legacy-prefix --execute --max-keys=500
|
||||
|
||||
# E presence (repeat)
|
||||
php artisan skinbase:prune-online-visitor-index --dry-run --max-batches=5
|
||||
php artisan skinbase:prune-online-visitor-index --execute --max-batches=200
|
||||
# then set ONLINE_VISITOR_INDEX_PRUNE_ENABLED=true and config:cache
|
||||
|
||||
# F collections (optional later)
|
||||
php artisan skinbase:redis-cleanup-orphans collections --dry-run
|
||||
php artisan skinbase:redis-cleanup-orphans collections --execute
|
||||
```
|
||||
|
||||
Abort: skip `--execute`. UNLINK is async; cannot restore payloads. Presence SREM is idempotent; live visitors re-SADD on next request.
|
||||
|
||||
---
|
||||
|
||||
## 14. Files
|
||||
|
||||
```text
|
||||
app/Support/Redis/OrphanQueueCleanup.php
|
||||
app/Services/Traffic/PresenceIndexPruner.php
|
||||
app/Console/Commands/RedisCleanupOrphansCommand.php
|
||||
app/Console/Commands/PruneOnlineVisitorIndexCommand.php
|
||||
app/Console/Commands/RedisCleanupLegacyPrefixCommand.php
|
||||
config/traffic.php
|
||||
routes/console.php
|
||||
.env.example
|
||||
tests/Feature/Redis/OrphanQueueCleanupTest.php
|
||||
tests/Unit/Traffic/PresenceIndexPrunerTest.php
|
||||
tests/Unit/Redis/LegacyPrefixFilterTest.php
|
||||
docs/optimization-m8-redis-orphan-cleanup.md
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 15. Risks
|
||||
|
||||
- Sampled classes (400 LINDEX) might miss a rare job type; `--force` required then.
|
||||
- `queues:*:notify` UNLINK with the list is required for Laravel blocking pop leftovers.
|
||||
- Presence first pass is slow; enable scheduler only after a few manual batches.
|
||||
- Predis fatals: do **not** LRANGE forum-moderation in tinker/Horizon UI until UNLINK.
|
||||
- Collections UNLINK loses queued refresh intent; DB is source of truth.
|
||||
|
||||
No FLUSH, no DEL of huge keys from HTTP, no Horizon workers, no Redis restart.
|
||||
@@ -0,0 +1,212 @@
|
||||
# M9 — Scheduler, Cron Overlap & Background Job Audit
|
||||
|
||||
```text
|
||||
STATUS: COMPLETE (audit + small local schedule hygiene)
|
||||
PRODUCTION: unchanged (no cron/systemd/Supervisor edits)
|
||||
WHEN: 2026-08-23 20:23–20:25 CEST
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 1. Scheduler execution mechanism
|
||||
|
||||
Laravel schedule is defined only in `routes/console.php` (`bootstrap/app.php` `commands:`). `Kernel::schedule()` is empty to avoid double registration.
|
||||
|
||||
**How `schedule:run` is invoked:** empirically **running** (live `schedule:finish` wrapper around `leaderboards:refresh`; `health:tick` Redis key present). User crontab empty. `/etc/cron.d` has **no** `artisan schedule:run`. systemd has **no** schedule timer. **Inferred: root crontab** (`sudo crontab -l` not readable by this user). Do not change.
|
||||
|
||||
Supervisor runs **Horizon + SSR only**, not `schedule:work`.
|
||||
|
||||
---
|
||||
|
||||
## 2. Duplicate-runner check
|
||||
|
||||
| Check | Result |
|
||||
| ----- | ------ |
|
||||
| `schedule:work` processes | **none** |
|
||||
| Second `schedule:run` host | single host `server3` |
|
||||
| Duplicate cron.d artisan | **none** |
|
||||
| Kernel + console.php double register | Kernel schedule empty; `schedule:list` unique |
|
||||
| Stale schedule:work | **none** |
|
||||
|
||||
**NO ISSUE** for duplicate runners.
|
||||
|
||||
---
|
||||
|
||||
## 3. Laravel schedule inventory (production `schedule:list`)
|
||||
|
||||
Gates: collections dispatch **false**, forum AI/sec **false**, presence prune **true**, security-report **true**.
|
||||
|
||||
Forum async and collections:dispatch-maintenance **absent** from list. Presence prune **present**.
|
||||
|
||||
Every-minute: posts/artworks/news/nova-cards publish, health:tick.
|
||||
|
||||
Heavy hourly: metrics :02, rankings :07/:37, heat :09/:24/:39/:54, rank-build :15, forum:scan-posts :17, leaderboards :21 (mutex held, **>2 min** observed), search-reconcile :28, presence prune :33, horizon snapshot :45.
|
||||
|
||||
Nightly rec: tags 02:00, behavior 02:15, hybrid 02:30 (queue jobs; M1/M7.1).
|
||||
|
||||
Sitemaps: generate 10:30 and 22:30; publish `8 */6`; validate 04:45; cleanup 05:00.
|
||||
|
||||
Full table: production `schedule:list` output captured 20:23 CEST (44 events). Local adds overlap names on analytics/uploads; flush minutes **18 instead of 21**.
|
||||
|
||||
`uploads:cleanup` previously had **no** `withoutOverlapping` (fixed locally). Four `analytics:aggregate-*` same (fixed locally).
|
||||
|
||||
Default overlap expiry: Laravel **1440 minutes** except prune-metric **120**, publish-scheduled **2**, presence prune **25**.
|
||||
|
||||
---
|
||||
|
||||
## 4. System cron/timers relevant to Skinbase
|
||||
|
||||
| Source | What |
|
||||
| ------ | ---- |
|
||||
| `/etc/cron.d/skinbase-public-index-integrity` | `*/5` root integrity script |
|
||||
| `/etc/cron.d/php` | sessionclean 09,39 (systemd timer wins) |
|
||||
| certbot.timer | ~22:26 |
|
||||
| logrotate.timer | ~00:02 |
|
||||
| apt-daily.timer | ~02:12 |
|
||||
| lynis, crowdsec-hub, dpkg-db-backup | ~00:00–00:26 |
|
||||
| ClamAV | daemon, not a Skinbase timer |
|
||||
|
||||
---
|
||||
|
||||
## 5–6. Timeline / collisions
|
||||
|
||||
| When | Tasks | Rank |
|
||||
| ---- | ----- | ---- |
|
||||
| **:21** | `flush-redis-stats` **and** `leaderboards:refresh` (refresh held mutex **>2 min**) | **MEDIUM** — local: flush moved to **:18** |
|
||||
| :33 | `collections:sync-lifecycle` (DB) + presence prune (Redis) | **LOW** — different resources |
|
||||
| :15 | rank-build-lists + homepage warm | **LOW** |
|
||||
| 02:00–02:30 | rec tags→behavior→hybrid on **default queue**; apt-daily ~02:12 | **MEDIUM** (queue contention, not schedule mutex). Offsets are guessed, not a completion barrier. Do not redesign in M9. |
|
||||
| 03:00–04:45 | uploads, analytics, reset-windowed, enhance, academy prune, metric prune, sitemap validate | **LOW–MEDIUM** overnight cluster, sequential minutes |
|
||||
| 03:30 | reset-windowed-stats-24h **and** enhance:cleanup | **LOW** (own mutexes, enhance background) |
|
||||
| every minute | 5 commands | **NO ISSUE** if each is cheap |
|
||||
|
||||
---
|
||||
|
||||
## 7. Measured runtimes
|
||||
|
||||
| Job | Evidence |
|
||||
| --- | -------- |
|
||||
| leaderboards:refresh | process **etime 2m23s** still running; schedule mutex listed |
|
||||
| IndexUserJob / MakeSearchable | Horizon 16–108 **ms** |
|
||||
| sitemaps generate | prior ~65s; `public/sitemaps/sitemap.xml` mtime 14:51 (release/publish, not 10:30) |
|
||||
| presence prune | first scheduled run **~81k** members (M8 verify); hourlyAt(33) keep |
|
||||
| rec nightly | historical MaxAttempts 02:18 / 02:31 (M1; mitigated in code, not schedule) |
|
||||
| health:tick | Redis SETEX; key `1787509506` present |
|
||||
|
||||
Did not run expensive commands for benchmarks.
|
||||
|
||||
---
|
||||
|
||||
## 8. withoutOverlapping
|
||||
|
||||
Missing (fixed locally): `uploads:cleanup`, four analytics aggregates.
|
||||
|
||||
TTL 1440m default: crash suppresses a daily job up to 24h — acceptable for daily. Presence **25m** matches catch-up batches. Metric prune **120m**. Publish **2m**.
|
||||
|
||||
leaderboards uses default 1440 while runtime ~2–3 min — OK.
|
||||
|
||||
---
|
||||
|
||||
## 9. runInBackground
|
||||
|
||||
Used for publish, rankings, metrics, heat, sitemaps, forum scan-posts, presence prune, leaderboards, etc.
|
||||
|
||||
Observed: `leaderboards:refresh` spawned via `schedule:finish` and stdout to `/dev/null`. Failures may not fail `schedule:run`. **Do not remove globally.** Mutex still applied (Has Mutex on list).
|
||||
|
||||
---
|
||||
|
||||
## 10. Every-minute
|
||||
|
||||
`health:tick`: one Redis SETEX, TTL 300. **NO CHANGE.**
|
||||
|
||||
Publish-scheduled ×3 + posts: **withoutOverlapping(2)** / default. Keep.
|
||||
|
||||
`posts:warm-trending` odd minutes: overlap protected.
|
||||
|
||||
---
|
||||
|
||||
## 11. Recommendation
|
||||
|
||||
02:00 tags, 02:15 behavior, 02:30 hybrid, everyFourHours pairs at **:00**. M1/M7.1 job code intact. **No completion wait** between stages — hybrid may run while tag batches still on default. Historical failures were retry_after (M1), not this offset. **No architecture change.**
|
||||
|
||||
---
|
||||
|
||||
## 12. Ranking / leaderboards
|
||||
|
||||
`nova:recalculate-rankings` :07/:37 background. `rank-build-lists` :15 (dispatches unique scopes, 6h lock). `leaderboards:refresh` :21 **DB-heavy, >2 min**. Duplicate scope stacking addressed in M7.1. **NO further frequency change.**
|
||||
|
||||
---
|
||||
|
||||
## 13. Metrics
|
||||
|
||||
Snapshot :02 (~50k upserts). Prune **04:25** overlap 120 — keep (outside :02). Studio 30d is HTTP not schedule.
|
||||
|
||||
---
|
||||
|
||||
## 14. Presence
|
||||
|
||||
Enabled on production. hourlyAt(33), overlap 25. Redis-heavy; :33 vs lifecycle DB. **Keep. Do not accelerate.**
|
||||
|
||||
---
|
||||
|
||||
## 15. Forum / collections
|
||||
|
||||
`forum:scan-posts --limit=250` hourly :17, in-process, overlap, background. Async forum **absent**. Collections dispatch **absent**. `collections:sync-lifecycle` every 10 min (3,13,23,33,43,53) — invite/lifecycle SQL, **not** the orphan queue. **NO CHANGE** to frequency without row-count evidence.
|
||||
|
||||
---
|
||||
|
||||
## 16. Sitemap
|
||||
|
||||
generate 10:30/22:30 background overlap. publish every 6h at :08. validate 04:45. cleanup 05:00. Two daily generates remain justified (fresh shards). **Do not regress M3.**
|
||||
|
||||
---
|
||||
|
||||
## 17. Horizon snapshot
|
||||
|
||||
hourlyAt(45). Current prefix `skinbase_horizon:`. Horizon log shows live job timings — metrics path works. **NO CHANGE.**
|
||||
|
||||
---
|
||||
|
||||
## 18. OS maintenance
|
||||
|
||||
apt-daily ~02:12 vs rec 02:00–02:30 (**MEDIUM** IO). logrotate ~00:02. Do not edit OS timers.
|
||||
|
||||
---
|
||||
|
||||
## 19. Deploy
|
||||
|
||||
`sync.sh` → `deploy-production.sh`, restarts Horizon/SSR. Background `schedule:finish` children can outlive symlink (leaderboards pattern). Mutex in cache/Redis (not files found on disk). **No deploy script change** — no proven double scheduler on release switch.
|
||||
|
||||
---
|
||||
|
||||
## 20. Ranked findings
|
||||
|
||||
| Sev | Finding | Action |
|
||||
| --- | ------- | ------ |
|
||||
| MEDIUM | flush-redis-stats and leaderboards both :21; leaderboards >2 min | **Local:** flush **:18** |
|
||||
| LOW | uploads:cleanup / analytics aggregates lacked withoutOverlapping | **Local:** added |
|
||||
| MEDIUM | rec stages overlap on default queue by clock | Document only |
|
||||
| LOW | :33 lifecycle + presence | Different stores; keep |
|
||||
| NO CHANGE | health:tick, presence hour, metric retention, forum/collections gates, sitemap times, Horizon snapshot, worker counts |
|
||||
|
||||
---
|
||||
|
||||
## 21. Recommended vs current (only implemented)
|
||||
|
||||
| Task | Current | Recommended |
|
||||
| ---- | ------- | ----------- |
|
||||
| flush-redis-stats | 1,11,**21**,31,41,51 | 1,11,**18**,31,41,51 |
|
||||
| uploads:cleanup | daily 03:00, no overlap | + `withoutOverlapping` + name |
|
||||
| analytics aggregates | daily, no overlap | + `withoutOverlapping` + names |
|
||||
|
||||
Rollback: revert those three edits in `routes/console.php`.
|
||||
|
||||
---
|
||||
|
||||
## 22–25. Files / tests / deploy
|
||||
|
||||
`routes/console.php`, `docs/optimization-m9-scheduler-overlap.md`, `tests/Feature/Scheduler/ScheduleInventoryTest.php`.
|
||||
|
||||
Deploy: ship `routes/console.php`; no cron edit; no Horizon restart required for schedule (next `schedule:run` loads new events). Rollback: revert file.
|
||||
|
||||
**Do not** increase workers/FPM/Redis/MySQL/swap or re-enable orphan queues.
|
||||
@@ -0,0 +1,607 @@
|
||||
# Skinbase.org — M0 Production Baseline & Runtime Audit
|
||||
|
||||
```text
|
||||
STATUS: COMPLETE
|
||||
```
|
||||
|
||||
Primary origin runtime unknowns were resolved via **read-only `ssh server3`**. Some items remain UNKNOWN (MySQL `slow.log` not readable by this user; root crontab not readable without sudo; `performance_schema` denied to `skinbase@localhost`; `iostat` not installed).
|
||||
|
||||
**Nothing was modified on server3.** No restarts, no cache clears, no SQL writes, no deploys.
|
||||
|
||||
Evidence classes used below:
|
||||
|
||||
```text
|
||||
LOCAL CODE EVIDENCE
|
||||
PRODUCTION SERVER EVIDENCE
|
||||
PUBLIC CLOUDFLARE/HTTP EVIDENCE
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Executive Summary
|
||||
|
||||
```text
|
||||
Production OS: Debian 13 (trixie), kernel 6.12.96+deb13-cloud-amd64, x86_64
|
||||
CPU: AMD EPYC, 8 cores / 8 threads (1 socket)
|
||||
RAM: 23 GiB; ~8.6 GiB available; SWAP 2.0 GiB **almost fully used**
|
||||
Storage: 400G virtio disk `/dev/sda1` ext4 ~185G/394G (49%); rotational=0 (SSD/NVMe-like)
|
||||
nginx: 1.26.3; HTTP/2 on; gzip+brotli in nginx.conf; vhost skinbase.org.conf
|
||||
PHP: 8.4.24 (FPM + CLI)
|
||||
PHP-FPM: pool `skinbase`, dynamic, max_children=14, memory_limit=256M, terminate=120s
|
||||
MySQL: 8.4.11-11 local; DB `skinbase` 2758 MB; buffer pool 10 GB
|
||||
Redis: 8.10.1 local 127.0.0.1:6379; 878 MB used / 2 GB max; allkeys-lru
|
||||
Meilisearch: local systemd; health available; RSS ~401 MB
|
||||
Queue manager: Supervisor → `php artisan horizon` (NOT standalone queue:work)
|
||||
Cache store: **redis** (PRODUCTION; not repo default database)
|
||||
Session store: **redis**
|
||||
CDN: Cloudflare + cdn.skinbase.org RGW/S3
|
||||
Download acceleration: **OFF** (`download_accel_enabled=false`; no X-Accel location in vhost)
|
||||
Sitemap static serving: **ACTIVE** (nginx try_files + shared files). Index file stale (May 13).
|
||||
Debug mode: OFF
|
||||
Debugbar production: NOT PRESENT (packages-dev=0, no debugbar routes)
|
||||
```
|
||||
|
||||
```text
|
||||
PRODUCTION_APP_PATH=/opt/www/virtual/SkinbaseNova
|
||||
→ symlink → /opt/www/virtual/SkinbaseNova.releases/current
|
||||
→ /opt/www/virtual/SkinbaseNova.releases/releases/20260801-172803-5af95f65-dirty
|
||||
|
||||
Local commit: f52879edbb19ecc5d3e807bd4872016f3071dad7 (develop)
|
||||
Production commit: 5af95f65 (release name 5af95f65-dirty, 2026-08-01)
|
||||
Same revision: no
|
||||
Production git: not a git checkout (no .git in release)
|
||||
```
|
||||
|
||||
```text
|
||||
Active P0 findings:
|
||||
QUEUE-001 (90s vs 900s): MITIGATED — Horizon default workers timeout=960
|
||||
QUEUE rec-job failures: ACTIVE P0 — RecComputeSimilar* fail every night MaxAttemptsExceeded
|
||||
SITEMAP-001 live-build: NOT the current incident — static files exist
|
||||
SITEMAP index stale: ACTIVE P1/P0-SEO — sitemap.xml mtime 2026-05-13 while shards refresh daily
|
||||
STAT-001: NOT currently P0 at 2666 views/24h; design still sync writes (P1)
|
||||
DB-001: indexes PRESENT; residual UNKNOWN without slow.log. Snapshot table 1.85 GB is the dominant store.
|
||||
|
||||
Active P1 findings: swap exhaustion; mail queue unconsumed (LLEN=3); upload 20M vs app 50M; Sentry traces 100%; FPM 10s slowlog volume; CF HTML DYNAMIC; asset TTL 30d not 1y immutable
|
||||
|
||||
Major current bottleneck: MySQL row-read volume + 1.85 GB hourly snapshots; nightly rec jobs failing; memory pressure (swap full)
|
||||
|
||||
Largest remaining unknown: contents of /var/log/mysql/slow.log (permission)
|
||||
|
||||
M1 readiness: YES for planning. Do not implement in M0.
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 1. Repository State (LOCAL)
|
||||
|
||||
```text
|
||||
branch: develop
|
||||
HEAD: f52879edbb19ecc5d3e807bd4872016f3071dad7
|
||||
commit: f52879ed Current state with latest updates
|
||||
tree: clean at start of this SSH pass (ahead of origin/develop by 1)
|
||||
```
|
||||
|
||||
No reset/stash/clean on local or production.
|
||||
|
||||
---
|
||||
|
||||
## 2. Production Host (PRODUCTION SERVER)
|
||||
|
||||
```text
|
||||
hostname: server3
|
||||
FQDN: server3.klevze.si
|
||||
OS: Debian GNU/Linux 13 (trixie) 13.6
|
||||
kernel: 6.12.96+deb13-cloud-amd64
|
||||
arch: x86_64
|
||||
uptime: 24 days, 16:58 (sampled 2026-08-23 13:43 CEST)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 3. Hardware Baseline (PRODUCTION SERVER)
|
||||
|
||||
```text
|
||||
CPU model: AMD EPYC Processor (with IBPB)
|
||||
sockets: 1
|
||||
cores: 8
|
||||
threads: 8 (1 thread/core)
|
||||
RAM: 23 GiB
|
||||
swap: 2.0 GiB, ~2.0 GiB used, 2.7 MiB free
|
||||
disk: /dev/sda 400G → sda1 399.9G on /
|
||||
filesystem: ~394G, 185G used, 49%
|
||||
rotational: 0 → SSD/cloud SSD (not HDD)
|
||||
```
|
||||
|
||||
`dmidecode` not used (would need sudo).
|
||||
|
||||
---
|
||||
|
||||
## 4. Current System Load (PRODUCTION SERVER)
|
||||
|
||||
Short snapshot only (limitation: not peak-hour).
|
||||
|
||||
```text
|
||||
load average: 0.85, 2.47, 2.88
|
||||
vmstat 5s: idle ~76–83%, wa ~0–1%, r=3–6
|
||||
CPU notables: mysqld ~20%, containerd/dockerd ~45% each (other workloads on same host), redis ~5%
|
||||
RAM: 14 GiB used + 7 GiB cache; 8.6 GiB available
|
||||
swap: fully used
|
||||
```
|
||||
|
||||
Largest RSS:
|
||||
|
||||
| Process | RSS | Note |
|
||||
| ------- | --: | ---- |
|
||||
| mysqld | ~7.7 GiB | 32.7% |
|
||||
| python uvicorn × several | up to 1.3 GiB | vision/ML stack |
|
||||
| clamd | ~1.1 GiB | |
|
||||
| meilisearch | ~397–401 MiB | |
|
||||
| redis | ~164 MiB | |
|
||||
| php-fpm skinbase | ~95–100 MiB each | 6 workers sampled |
|
||||
| node ssr.js | ~122 MiB | |
|
||||
|
||||
Classification: **WATCH / CONCERNING** (swap exhausted; shared host with Docker/ML/Qdrant/Gitea/Crowdsec). Not a Skinbase-only box.
|
||||
|
||||
---
|
||||
|
||||
## 5. Process Inventory (PRODUCTION SERVER)
|
||||
|
||||
| Component | Reality |
|
||||
| --------- | ------- |
|
||||
| Horizon | **YES** — Supervisor `skinbase-horizon`, pid running ~1d |
|
||||
| standalone `queue:work` | **NO** |
|
||||
| Supervisor | **YES** — horizon + ssr |
|
||||
| systemd queue workers | **NO** |
|
||||
| Inertia SSR | **YES** — `node bootstrap/ssr/ssr.js` as www-data |
|
||||
| Reverb | **YES** — systemd `skinbase-reverb.service`, 127.0.0.1:8080 |
|
||||
| Meilisearch | **YES** — local systemd |
|
||||
| Redis | **local** 127.0.0.1:6379 |
|
||||
| MySQL | **local** `/usr/sbin/mysqld` |
|
||||
| PHP-FPM | php8.4, pool `skinbase` + unused pool `www` |
|
||||
| nginx | master + 8 workers, www-data |
|
||||
| Docker/ML | uvicorn:8000, qdrant, gitea — **co-tenant** |
|
||||
|
||||
---
|
||||
|
||||
## 6. nginx (PRODUCTION SERVER)
|
||||
|
||||
```text
|
||||
version: nginx/1.26.3
|
||||
vhost: /etc/nginx/sites-enabled/skinbase.org.conf
|
||||
root: /opt/www/virtual/SkinbaseNova/public
|
||||
listen: 80 → 301 HTTPS; 443 ssl; http2 on
|
||||
PHP: unix:/run/php/php8.4-fpm-skinbase.sock
|
||||
client_max_body_size: 24m
|
||||
gzip: on (nginx.conf)
|
||||
brotli: on (nginx.conf)
|
||||
real IP: /etc/nginx/conf.d/00-cloudflare-realip.conf ACTIVE
|
||||
search: location = /search + conf.d/13-skinbase-search-protection.conf (20r/m, abusive query map)
|
||||
HTTP/3: not in this vhost (Cloudflare alt-svc h3 is edge-only)
|
||||
```
|
||||
|
||||
Snippet vs live:
|
||||
|
||||
| Snippet | Status |
|
||||
| ------- | ------ |
|
||||
| sitemaps.conf (`max-age=21600`, try_files) | **ACTIVE** (inlined in vhost) |
|
||||
| static-cache.conf 1y immutable `/build/assets` | **NOT ACTIVE** — generic `expires 30d` for css/js/images |
|
||||
| download-accel.conf | **NOT ACTIVE** |
|
||||
| search-rate-limit.conf (repo) | **EQUIVALENT ACTIVE** as `13-skinbase-search-protection.conf` |
|
||||
| upstream-error-pages.conf | **NOT seen** in vhost |
|
||||
|
||||
**Bug/risk:** sitemap locations use `try_files $uri @php` but **no `location @php`** exists in the vhost. Missing files will not fall through to Laravel as the snippet comments claim.
|
||||
|
||||
No `nginx -T` via sudo (password required). Readable vhost was enough.
|
||||
|
||||
---
|
||||
|
||||
## 7–9. HTTP / CDN (PUBLIC CLOUDFLARE/HTTP + PRODUCTION)
|
||||
|
||||
Unchanged from edge probes, now correlated with origin:
|
||||
|
||||
- Homepage Cache-Control matches Laravel guest headers; CF **DYNAMIC**
|
||||
- Build assets CF **HIT**, origin `expires 30d`
|
||||
- CDN webp `max-age=31536000` HIT + Polish
|
||||
- Brotli **CONFIRMED** at edge; origin brotli **on**
|
||||
- HTTP/2 **CONFIRMED** at origin (`http2 on`); this Windows curl was HTTP/1.1 to CF
|
||||
- HTTP/3 **offered by CF** (`alt-svc`), not configured on origin vhost
|
||||
|
||||
---
|
||||
|
||||
## 10–12. PHP / OPcache / FPM (PRODUCTION SERVER)
|
||||
|
||||
```text
|
||||
PHP: 8.4.24
|
||||
FPM master: /etc/php/8.4/fpm/php-fpm.conf
|
||||
pool: [skinbase] user=skinbase
|
||||
listen: /run/php/php8.4-fpm-skinbase.sock
|
||||
pm: dynamic
|
||||
max_children: 14
|
||||
start_servers: 4
|
||||
min/max spare: 2 / 6
|
||||
max_requests: 500
|
||||
request_terminate_timeout: 120s
|
||||
slowlog: 10s → /var/log/php8.4-fpm-skinbase-slow.log (23 MB, 139556 lines)
|
||||
memory_limit: 256M (pool)
|
||||
max_execution: 60s (pool)
|
||||
upload/post: 20M / 24M (pool) ← app allows 50M images / 200M archives
|
||||
CLI upload_max: 2M (irrelevant to FPM)
|
||||
extensions: opcache, redis, pcntl, posix, intl, mbstring, curl, gd, pdo_mysql
|
||||
imagick: NOT in php -m
|
||||
opcache (FPM php.ini): memory 256MB, interned 16, max_files 20000, validate_timestamps=1, revalidate_freq=2, jit=off
|
||||
opcache runtime hit rate: UNKNOWN (no status scrape of /fpm-status from remote; localhost-only)
|
||||
```
|
||||
|
||||
Workers sampled: **6** skinbase processes, **avg RSS ~95 MB**, total ~570 MB.
|
||||
|
||||
Theoretical ceiling: `14 × 256 MB = 3.6 GB` if every worker hit `memory_limit`. Typical: `14 × 95 MB ≈ 1.3 GB`.
|
||||
|
||||
Available RAM 8.6 GiB **but swap already full** — do not raise max_children in M0.
|
||||
|
||||
---
|
||||
|
||||
## 13–15. Laravel Production Configuration (PRODUCTION SERVER)
|
||||
|
||||
`php artisan about` + tinker `config()` (no `.env` dump):
|
||||
|
||||
```text
|
||||
APP_ENV=production
|
||||
APP_DEBUG=false
|
||||
CACHE_STORE=redis
|
||||
SESSION_DRIVER=redis
|
||||
QUEUE_CONNECTION=redis
|
||||
SCOUT_DRIVER=meilisearch
|
||||
FILESYSTEM_DISK=local
|
||||
UPLOAD_QUEUE_DERIVATIVES=false
|
||||
DOWNLOAD_ACCEL_ENABLED=false
|
||||
DOWNLOAD_ACCEL_PATH=/internal/originals (path set, flag off, nginx location missing)
|
||||
SITEMAPS_BUILD_ON_REQUEST=true
|
||||
SITEMAPS_FALLBACK_TO_LIVE_BUILD=true
|
||||
SITEMAPS_PREGENERATED_ENABLED=true
|
||||
SITEMAPS_PREFER_PUBLISHED_RELEASE=true
|
||||
SITEMAPS_STATIC_PUBLISH_ENABLED=true
|
||||
REDIS_CLIENT=predis
|
||||
homepage.cache_store=homepage
|
||||
homepage.guest_payload_ttl_seconds=1800
|
||||
recommendations.queue=default
|
||||
vision.queue=default
|
||||
discovery.queue=default
|
||||
broadcasting=reverb
|
||||
horizon.path=horizon
|
||||
Sentry enabled, sample rate errors 100%, performance 100%
|
||||
```
|
||||
|
||||
Laravel caches: **config, events, routes, views ALL CACHED**.
|
||||
|
||||
Debugbar/Telescope/Clockwork routes: **none**. Composer `packages-dev=0` → **--no-dev install**.
|
||||
|
||||
---
|
||||
|
||||
## 16–22. MySQL (PRODUCTION SERVER)
|
||||
|
||||
```text
|
||||
location: local mysqld
|
||||
version: 8.4.11-11
|
||||
database: skinbase
|
||||
size: 2758.2 MB (data 1230.4 + indexes 1527.8)
|
||||
buffer pool: 10 GB, 10 instances
|
||||
max_conn: 120 (Max_used 25)
|
||||
slow_query_log: ON, long_query_time=0.5s, file /var/log/mysql/slow.log (not readable here)
|
||||
Slow_queries: 11342 in Uptime 97330s (~27h) ≈ 0.12/s
|
||||
Threads: connected 3, running 2
|
||||
InnoDB hit: 1 - 103791/18420885623 ≈ 99.999%
|
||||
Rows read: 9.70e9 in ~27h (high — job/scan load)
|
||||
tmp tables: 455567 created, 21 on disk
|
||||
```
|
||||
|
||||
Largest tables (information_schema estimates):
|
||||
|
||||
| Table | ~rows | total MB |
|
||||
| ----- | -----: | -------: |
|
||||
| artwork_metric_snapshots_hourly | 8.62M | **1855.5** |
|
||||
| artwork_downloads | 579k | 167.3 |
|
||||
| forum_bot_logs | 267k | 141.8 |
|
||||
| artworks | 50k | 94.2 |
|
||||
| user_activities | 240k | 69.3 |
|
||||
| artwork_comments | 172k | 62.2 |
|
||||
| artwork_view_events | 161k | 25.4 |
|
||||
| rank_artwork_scores | 50k | 22.4 |
|
||||
| rec_item_pairs | 74k | 16.0 |
|
||||
| artwork_stats | 49k | 15.2 |
|
||||
|
||||
`cache`/`sessions` tables nearly empty (drivers are Redis). `jobs` table empty (Redis queues).
|
||||
|
||||
### Indexes (batch1) — PRODUCTION
|
||||
|
||||
| Index | Status |
|
||||
| ----- | ------ |
|
||||
| artworks.idx_public_approved_published_id | **PRESENT** |
|
||||
| artworks.idx_public_approved_user_id | **PRESENT** |
|
||||
| artworks FULLTEXT title+description | **PRESENT** |
|
||||
| snapshots.idx_bucket_artwork | **PRESENT** |
|
||||
| rank_artwork_scores idx_mv_trending/new_hot/best | **PRESENT** |
|
||||
| tags.artworks_count + idx_tags_artworks_count | **PRESENT** |
|
||||
|
||||
### DB-001 re-evaluation
|
||||
|
||||
```text
|
||||
Classification: UNKNOWN as “still 78%”, NOT “unindexed public scans”
|
||||
```
|
||||
|
||||
April indexes **are deployed**. Dominant storage is hourly snapshots ≈ 50k artworks × 24h × 7d. Heat/ranking jobs that join this table can still dominate `Innodb_rows_read` even with `idx_bucket_artwork`. Cannot confirm fingerprints without `slow.log` / performance_schema (denied).
|
||||
|
||||
Do not add more indexes in M0.
|
||||
|
||||
---
|
||||
|
||||
## 23–25. Redis / Cache / Sessions (PRODUCTION SERVER)
|
||||
|
||||
```text
|
||||
redis_version: 8.10.1
|
||||
used: 877.55 MB (peak 917 MB)
|
||||
maxmemory: 2.00 G, policy allkeys-lru
|
||||
evicted_keys: 0
|
||||
clients: 20
|
||||
ops/sec: 44 (instant)
|
||||
keyspace hits/misses: 1.40M / 1.30M (~52% hit)
|
||||
db0: 1932 keys; db1: 6641; db2: 55176 (all with TTL)
|
||||
```
|
||||
|
||||
| System | Redis in production? |
|
||||
| ------ | -------------------- |
|
||||
| Cache default | **YES** |
|
||||
| Homepage store `homepage` | configured (failover redis→database) |
|
||||
| Sessions | **YES** |
|
||||
| Queues / Horizon | **YES** |
|
||||
| Presence | code uses Redis (not separately counted) |
|
||||
| Stats deltas | flush command scheduled; view path still MySQL sync |
|
||||
|
||||
CACHE-001 from the code audit (**database cache default**) is **NOT ACTIVE in production**.
|
||||
|
||||
---
|
||||
|
||||
## 26–32. Horizon / Queues / QUEUE-001 (PRODUCTION SERVER)
|
||||
|
||||
Horizon **running**. Supervisors:
|
||||
|
||||
| Supervisor | queues | timeout | maxProcesses |
|
||||
| ---------- | ------ | ------: | -----------: |
|
||||
| supervisor-default | search, default | **960** | 5 |
|
||||
| supervisor-messaging | broadcasts, notifications | **90** | 3 |
|
||||
|
||||
`recommendations.queue` / `vision.queue` / `discovery.queue` = **`default`** → consumed by 960s workers.
|
||||
|
||||
Standalone supervisor `queue:work --timeout=90` from `deploy/supervisor/skinbase-queue.conf` is **NOT the production worker**.
|
||||
|
||||
Failed jobs: **329** rows (`queue:failed` listed many; table_rows estimate 159). Dominant:
|
||||
|
||||
```text
|
||||
RecComputeSimilarByBehaviorJob ~02:18 daily MaxAttemptsExceededException
|
||||
RecComputeSimilarHybridJob ~02:31 daily MaxAttemptsExceededException
|
||||
```
|
||||
|
||||
Horizon worker `tries=1`, job `$timeout=900`, worker timeout 960. Failures are **not** explained by the old 90s worker. Likely 900s job timeout, 128 MB worker memory, or lock/overlap. **Do not retry in M0.**
|
||||
|
||||
Queue LLEN (Redis): all listed queues 0 except **`queues:mail` = 3**. Horizon does **not** listen to `mail`. Mail jobs can stall.
|
||||
|
||||
```text
|
||||
QUEUE-001 (90 vs 900): MITIGATED
|
||||
Nightly rec job failures: ACTIVE P0 (related, different mechanism)
|
||||
mail queue unconsumed: P1
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 31–32. Scheduler (PRODUCTION SERVER)
|
||||
|
||||
`php artisan schedule:list` matches `routes/console.php` (generate sitemaps 10:30/22:30, rec jobs 02:00–02:30, etc.).
|
||||
|
||||
**How `schedule:run` is invoked:** user crontab empty; sudo crontab **not readable**; no systemd timer named laravel/schedule. **Empirically running** (sitemap shards 10:30 today; rec jobs fail 02:18/02:31 daily). Likely root crontab. Do not change.
|
||||
|
||||
02:00–05:00 window (code + evidence it fires): rec jobs, ranking/heat, analytics, prune snapshots, sitemap validate — plus nightly rec **failures**.
|
||||
|
||||
---
|
||||
|
||||
## 33. Meilisearch (PRODUCTION SERVER)
|
||||
|
||||
```text
|
||||
process: /usr/local/bin/meilisearch --config-file-path /etc/meilisearch.toml
|
||||
health: {"status":"available"}
|
||||
version/stats: 401 without key (not printed)
|
||||
RSS: ~401 MB
|
||||
```
|
||||
|
||||
No reindex.
|
||||
|
||||
---
|
||||
|
||||
## 34–36. Storage / Downloads / Sitemaps
|
||||
|
||||
- `FILESYSTEM_DISK=local`; CDN for public derivatives.
|
||||
- Download accel **off**; originals would stream via PHP if used.
|
||||
- Sitemaps live under shared `.../shared/public/sitemaps/` (20M, 27 xml files).
|
||||
- **Shards + academy + users + forum-threads mtime 2026-08-23 10:30–10:31**
|
||||
- **`sitemap.xml` mtime 2026-05-13 21:39, 1906 bytes** — generate/publish does **not** update the index file.
|
||||
- May index omits academy-* families that now exist on disk.
|
||||
|
||||
SITEMAP-001 live-build-on-missing: files exist, so crawlers are not building XML in PHP for the index. `@php` named location missing anyway.
|
||||
|
||||
```text
|
||||
SITEMAP-001 original (PHP live-build stampede): NOT ACTIVE right now
|
||||
Stale sitemap index vs daily shards: ACTIVE (SEO/ops)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 37–38. Views / Downloads
|
||||
|
||||
```text
|
||||
artwork_view_events last 1h: 72
|
||||
artwork_view_events last 24h: 2666
|
||||
artwork_downloads last 24h: 4718
|
||||
view_events table: ~161k rows, 25.4 MB
|
||||
downloads table: ~579k rows, 167 MB
|
||||
```
|
||||
|
||||
STAT-001 architecture (sync INSERT+UPDATE, `defer=false`) is **still the code on this release**. At **~0.03 views/s** it is not the current DB bottleneck. Classify **P1** (architecture / growth), not active P0 load.
|
||||
|
||||
Downloads were not fetched (would write).
|
||||
|
||||
---
|
||||
|
||||
## 39–47. Latency / Traffic / Resources
|
||||
|
||||
Origin access-log QPS not aggregated (would include IPs — skipped). Edge n=1 TTFBs from earlier remain **PUBLIC** evidence, not origin APM.
|
||||
|
||||
FPM slowlog: 139k historical lines, last file mtime Aug 23 05:05, timeout 10s. Scripts are `index.php` only (no URI dump here).
|
||||
|
||||
Resource baseline: see §4. **Swap full = CONCERNING.**
|
||||
|
||||
---
|
||||
|
||||
## 48. Lighthouse / CWV
|
||||
|
||||
Not re-run. Prior lab file is skinbase.top 2026-03-23 — do not use as this baseline.
|
||||
|
||||
---
|
||||
|
||||
## 49. Observability
|
||||
|
||||
- Sentry **on**, 100% error **and** performance sample (P1 cost).
|
||||
- Netdata on host.
|
||||
- Horizon dashboard path `horizon` (auth not verified; do not probe).
|
||||
- `/fpm-status` localhost only.
|
||||
- `/stats/` basic-auth (not accessed).
|
||||
- MySQL slow.log exists but unreadable to this SSH user.
|
||||
|
||||
---
|
||||
|
||||
## 50. Re-evaluated P0 Findings
|
||||
|
||||
| ID | Verdict | Why |
|
||||
| -- | ------- | --- |
|
||||
| QUEUE-001 90s worker vs 900s job | **MITIGATED** | Horizon default timeout **960**; rec queues aliased to `default` |
|
||||
| RecCompute* nightly MaxAttemptsExceeded | **ACTIVE P0** | every night 02:18/02:31 on `redis@default`; 329 failed jobs |
|
||||
| STAT-001 sync view writes | **P1** at current 2666/day | still sync in config/code; not the load leader |
|
||||
| SITEMAP-001 request-time build | **NOT ACTIVE** | static files present; `@php` missing |
|
||||
| Stale sitemap.xml (May vs Aug shards) | **ACTIVE P1** (SEO) | generate writes children, not index |
|
||||
| DB-001 unindexed aggregates | **MITIGATED for missing indexes** | batch1 **PRESENT**. Residual scan cost **UNKNOWN**; snapshots 1.85 GB |
|
||||
|
||||
---
|
||||
|
||||
## 51. Re-evaluated P1 Findings
|
||||
|
||||
- Production **Redis** cache/session (repo default database is wrong for prod).
|
||||
- Upload 20M/24M FPM vs app 50M/200M.
|
||||
- `UPLOAD_QUEUE_DERIVATIVES=false` — sync image work on publish.
|
||||
- `mail` Redis queue not consumed by Horizon.
|
||||
- Sentry performance sampling 100%.
|
||||
- CF HTML DYNAMIC despite public Cache-Control.
|
||||
- Asset TTL 30d not 1y immutable.
|
||||
- Swap 2G full; shared ML/Docker host.
|
||||
- FPM slowlog 10s historically large.
|
||||
- Local git **newer** than production release (f52879ed vs 5af95f65).
|
||||
- Session cookies on `/search` and 404 (code-consistent).
|
||||
- Duplicate security headers (app middleware + nginx snippet).
|
||||
|
||||
---
|
||||
|
||||
## 52. Baseline Metric Table
|
||||
|
||||
| Metric | Value | Source |
|
||||
| ------ | ----- | ------ |
|
||||
| DB size | 2758 MB | information_schema |
|
||||
| Snapshots table | 1855 MB / ~8.6M rows | information_schema |
|
||||
| Artworks | ~50k / 94 MB | information_schema |
|
||||
| Views 24h | 2666 | count on 161k table |
|
||||
| Downloads 24h | 4718 | count |
|
||||
| Slow_queries / 27h | 11342 | SHOW STATUS |
|
||||
| InnoDB pool hit | ~100% | STATUS |
|
||||
| Redis used | 878 MB / 2 GB | INFO |
|
||||
| Redis hit ratio | ~52% | INFO |
|
||||
| FPM workers | 6 / max 14, ~95 MB RSS | ps |
|
||||
| Horizon | running, timeout 960/90 | ps + artisan |
|
||||
| Failed jobs | 329 | DB |
|
||||
| Homepage TTFB (edge) | 146 ms n=1 | PUBLIC |
|
||||
| Production release | 20260801-172803-5af95f65-dirty | filesystem |
|
||||
|
||||
---
|
||||
|
||||
## 53. Production Risk Matrix
|
||||
|
||||
| Risk | Evidence | Impact | Confidence |
|
||||
| ---- | -------- | ------ | ---------- |
|
||||
| Nightly rec jobs never succeed | failed_jobs every day | stale similar-art | high |
|
||||
| Hourly snapshot table 1.85 GB | information_schema | heat/rank I/O | high |
|
||||
| Swap full + co-tenant ML | free/ps | latency spikes | high |
|
||||
| Stale sitemap index | mtime May 13 vs shards today | crawl/SEO | high |
|
||||
| mail queue not consumed | LLEN=3, Horizon queues | delayed mail | high |
|
||||
| Missing @php location | nginx vhost | missing shard → error not Laravel | high |
|
||||
| Cannot read slow.log | permissions | DB-001 residual | high |
|
||||
|
||||
---
|
||||
|
||||
## 54. Optimization Targets (do not implement now)
|
||||
|
||||
1. Rec job failures (timeout/memory/overlap) — after measuring why MaxAttemptsExceeded.
|
||||
2. Snapshot retention / heat query EXPLAIN once slow.log is readable.
|
||||
3. Publish `sitemap.xml` index with shards.
|
||||
4. Consume `mail` or stop using that queue name.
|
||||
5. Download X-Accel if originals still hit FPM.
|
||||
6. Align upload limits.
|
||||
7. Memory/swap / co-tenancy.
|
||||
8. Sentry sample rates.
|
||||
9. CF HTML cache policy (product decision).
|
||||
|
||||
---
|
||||
|
||||
## 55. M1–M4 Readiness
|
||||
|
||||
```text
|
||||
M1 planning: READY (production drivers known; P0 recast)
|
||||
M1 implementation: not this milestone
|
||||
Blockers for DB-001 closeout: mysql slow.log or GRANTs on performance_schema
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 56. Unknowns
|
||||
|
||||
- Root crontab contents (scheduler invocation path)
|
||||
- MySQL slow.log fingerprints **now**
|
||||
- OPcache hit rate (FPM status localhost-only)
|
||||
- Meilisearch document counts (auth)
|
||||
- Peak QPS / access-log (IPs omitted on purpose)
|
||||
- Why rec jobs MaxAttemptsExceeded (timeout vs memory vs exception) — do not dump payloads
|
||||
- HTTP/3 to origin (not configured)
|
||||
- iostat (command missing)
|
||||
|
||||
---
|
||||
|
||||
## 57. Repository Changes
|
||||
|
||||
```text
|
||||
Created earlier:
|
||||
- scripts/collect-production-baseline.sh
|
||||
|
||||
Updated:
|
||||
- docs/skinbase-production-baseline.md
|
||||
|
||||
Modified application:
|
||||
- none
|
||||
|
||||
Database changes:
|
||||
- none
|
||||
|
||||
Server configuration changes:
|
||||
- none
|
||||
|
||||
Service restarts:
|
||||
- none
|
||||
|
||||
Cache clears:
|
||||
- none
|
||||
|
||||
Queue operations:
|
||||
- none (failed jobs listed only, not retried)
|
||||
```
|
||||
Reference in New Issue
Block a user