Ship production optimization M1-M12.5A: queues, metrics, HTTP observability, and vector search reliability.

Keep similar-ai from tripping the global circuit on a lone URL 502, clamp Qdrant search to 100, and add Server-Timing plus slow-request logging. Studio shared props, Academy S3 exists caching, heat chunking, and Redis/scheduler hygiene stay in this rollout.
This commit is contained in:
2026-08-25 07:58:47 +02:00
parent f52879edbb
commit 8a80aae21e
114 changed files with 9293 additions and 284 deletions
+607
View File
@@ -0,0 +1,607 @@
# Skinbase.org — M0 Production Baseline & Runtime Audit
```text
STATUS: COMPLETE
```
Primary origin runtime unknowns were resolved via **read-only `ssh server3`**. Some items remain UNKNOWN (MySQL `slow.log` not readable by this user; root crontab not readable without sudo; `performance_schema` denied to `skinbase@localhost`; `iostat` not installed).
**Nothing was modified on server3.** No restarts, no cache clears, no SQL writes, no deploys.
Evidence classes used below:
```text
LOCAL CODE EVIDENCE
PRODUCTION SERVER EVIDENCE
PUBLIC CLOUDFLARE/HTTP EVIDENCE
```
---
## Executive Summary
```text
Production OS: Debian 13 (trixie), kernel 6.12.96+deb13-cloud-amd64, x86_64
CPU: AMD EPYC, 8 cores / 8 threads (1 socket)
RAM: 23 GiB; ~8.6 GiB available; SWAP 2.0 GiB **almost fully used**
Storage: 400G virtio disk `/dev/sda1` ext4 ~185G/394G (49%); rotational=0 (SSD/NVMe-like)
nginx: 1.26.3; HTTP/2 on; gzip+brotli in nginx.conf; vhost skinbase.org.conf
PHP: 8.4.24 (FPM + CLI)
PHP-FPM: pool `skinbase`, dynamic, max_children=14, memory_limit=256M, terminate=120s
MySQL: 8.4.11-11 local; DB `skinbase` 2758 MB; buffer pool 10 GB
Redis: 8.10.1 local 127.0.0.1:6379; 878 MB used / 2 GB max; allkeys-lru
Meilisearch: local systemd; health available; RSS ~401 MB
Queue manager: Supervisor → `php artisan horizon` (NOT standalone queue:work)
Cache store: **redis** (PRODUCTION; not repo default database)
Session store: **redis**
CDN: Cloudflare + cdn.skinbase.org RGW/S3
Download acceleration: **OFF** (`download_accel_enabled=false`; no X-Accel location in vhost)
Sitemap static serving: **ACTIVE** (nginx try_files + shared files). Index file stale (May 13).
Debug mode: OFF
Debugbar production: NOT PRESENT (packages-dev=0, no debugbar routes)
```
```text
PRODUCTION_APP_PATH=/opt/www/virtual/SkinbaseNova
→ symlink → /opt/www/virtual/SkinbaseNova.releases/current
→ /opt/www/virtual/SkinbaseNova.releases/releases/20260801-172803-5af95f65-dirty
Local commit: f52879edbb19ecc5d3e807bd4872016f3071dad7 (develop)
Production commit: 5af95f65 (release name 5af95f65-dirty, 2026-08-01)
Same revision: no
Production git: not a git checkout (no .git in release)
```
```text
Active P0 findings:
QUEUE-001 (90s vs 900s): MITIGATED — Horizon default workers timeout=960
QUEUE rec-job failures: ACTIVE P0 — RecComputeSimilar* fail every night MaxAttemptsExceeded
SITEMAP-001 live-build: NOT the current incident — static files exist
SITEMAP index stale: ACTIVE P1/P0-SEO — sitemap.xml mtime 2026-05-13 while shards refresh daily
STAT-001: NOT currently P0 at 2666 views/24h; design still sync writes (P1)
DB-001: indexes PRESENT; residual UNKNOWN without slow.log. Snapshot table 1.85 GB is the dominant store.
Active P1 findings: swap exhaustion; mail queue unconsumed (LLEN=3); upload 20M vs app 50M; Sentry traces 100%; FPM 10s slowlog volume; CF HTML DYNAMIC; asset TTL 30d not 1y immutable
Major current bottleneck: MySQL row-read volume + 1.85 GB hourly snapshots; nightly rec jobs failing; memory pressure (swap full)
Largest remaining unknown: contents of /var/log/mysql/slow.log (permission)
M1 readiness: YES for planning. Do not implement in M0.
```
---
## 1. Repository State (LOCAL)
```text
branch: develop
HEAD: f52879edbb19ecc5d3e807bd4872016f3071dad7
commit: f52879ed Current state with latest updates
tree: clean at start of this SSH pass (ahead of origin/develop by 1)
```
No reset/stash/clean on local or production.
---
## 2. Production Host (PRODUCTION SERVER)
```text
hostname: server3
FQDN: server3.klevze.si
OS: Debian GNU/Linux 13 (trixie) 13.6
kernel: 6.12.96+deb13-cloud-amd64
arch: x86_64
uptime: 24 days, 16:58 (sampled 2026-08-23 13:43 CEST)
```
---
## 3. Hardware Baseline (PRODUCTION SERVER)
```text
CPU model: AMD EPYC Processor (with IBPB)
sockets: 1
cores: 8
threads: 8 (1 thread/core)
RAM: 23 GiB
swap: 2.0 GiB, ~2.0 GiB used, 2.7 MiB free
disk: /dev/sda 400G → sda1 399.9G on /
filesystem: ~394G, 185G used, 49%
rotational: 0 → SSD/cloud SSD (not HDD)
```
`dmidecode` not used (would need sudo).
---
## 4. Current System Load (PRODUCTION SERVER)
Short snapshot only (limitation: not peak-hour).
```text
load average: 0.85, 2.47, 2.88
vmstat 5s: idle ~76–83%, wa ~0–1%, r=3–6
CPU notables: mysqld ~20%, containerd/dockerd ~45% each (other workloads on same host), redis ~5%
RAM: 14 GiB used + 7 GiB cache; 8.6 GiB available
swap: fully used
```
Largest RSS:
| Process | RSS | Note |
| ------- | --: | ---- |
| mysqld | ~7.7 GiB | 32.7% |
| python uvicorn × several | up to 1.3 GiB | vision/ML stack |
| clamd | ~1.1 GiB | |
| meilisearch | ~397–401 MiB | |
| redis | ~164 MiB | |
| php-fpm skinbase | ~95–100 MiB each | 6 workers sampled |
| node ssr.js | ~122 MiB | |
Classification: **WATCH / CONCERNING** (swap exhausted; shared host with Docker/ML/Qdrant/Gitea/Crowdsec). Not a Skinbase-only box.
---
## 5. Process Inventory (PRODUCTION SERVER)
| Component | Reality |
| --------- | ------- |
| Horizon | **YES** — Supervisor `skinbase-horizon`, pid running ~1d |
| standalone `queue:work` | **NO** |
| Supervisor | **YES** — horizon + ssr |
| systemd queue workers | **NO** |
| Inertia SSR | **YES** — `node bootstrap/ssr/ssr.js` as www-data |
| Reverb | **YES** — systemd `skinbase-reverb.service`, 127.0.0.1:8080 |
| Meilisearch | **YES** — local systemd |
| Redis | **local** 127.0.0.1:6379 |
| MySQL | **local** `/usr/sbin/mysqld` |
| PHP-FPM | php8.4, pool `skinbase` + unused pool `www` |
| nginx | master + 8 workers, www-data |
| Docker/ML | uvicorn:8000, qdrant, gitea — **co-tenant** |
---
## 6. nginx (PRODUCTION SERVER)
```text
version: nginx/1.26.3
vhost: /etc/nginx/sites-enabled/skinbase.org.conf
root: /opt/www/virtual/SkinbaseNova/public
listen: 80 → 301 HTTPS; 443 ssl; http2 on
PHP: unix:/run/php/php8.4-fpm-skinbase.sock
client_max_body_size: 24m
gzip: on (nginx.conf)
brotli: on (nginx.conf)
real IP: /etc/nginx/conf.d/00-cloudflare-realip.conf ACTIVE
search: location = /search + conf.d/13-skinbase-search-protection.conf (20r/m, abusive query map)
HTTP/3: not in this vhost (Cloudflare alt-svc h3 is edge-only)
```
Snippet vs live:
| Snippet | Status |
| ------- | ------ |
| sitemaps.conf (`max-age=21600`, try_files) | **ACTIVE** (inlined in vhost) |
| static-cache.conf 1y immutable `/build/assets` | **NOT ACTIVE** — generic `expires 30d` for css/js/images |
| download-accel.conf | **NOT ACTIVE** |
| search-rate-limit.conf (repo) | **EQUIVALENT ACTIVE** as `13-skinbase-search-protection.conf` |
| upstream-error-pages.conf | **NOT seen** in vhost |
**Bug/risk:** sitemap locations use `try_files $uri @php` but **no `location @php`** exists in the vhost. Missing files will not fall through to Laravel as the snippet comments claim.
No `nginx -T` via sudo (password required). Readable vhost was enough.
---
## 7–9. HTTP / CDN (PUBLIC CLOUDFLARE/HTTP + PRODUCTION)
Unchanged from edge probes, now correlated with origin:
- Homepage Cache-Control matches Laravel guest headers; CF **DYNAMIC**
- Build assets CF **HIT**, origin `expires 30d`
- CDN webp `max-age=31536000` HIT + Polish
- Brotli **CONFIRMED** at edge; origin brotli **on**
- HTTP/2 **CONFIRMED** at origin (`http2 on`); this Windows curl was HTTP/1.1 to CF
- HTTP/3 **offered by CF** (`alt-svc`), not configured on origin vhost
---
## 10–12. PHP / OPcache / FPM (PRODUCTION SERVER)
```text
PHP: 8.4.24
FPM master: /etc/php/8.4/fpm/php-fpm.conf
pool: [skinbase] user=skinbase
listen: /run/php/php8.4-fpm-skinbase.sock
pm: dynamic
max_children: 14
start_servers: 4
min/max spare: 2 / 6
max_requests: 500
request_terminate_timeout: 120s
slowlog: 10s → /var/log/php8.4-fpm-skinbase-slow.log (23 MB, 139556 lines)
memory_limit: 256M (pool)
max_execution: 60s (pool)
upload/post: 20M / 24M (pool) ← app allows 50M images / 200M archives
CLI upload_max: 2M (irrelevant to FPM)
extensions: opcache, redis, pcntl, posix, intl, mbstring, curl, gd, pdo_mysql
imagick: NOT in php -m
opcache (FPM php.ini): memory 256MB, interned 16, max_files 20000, validate_timestamps=1, revalidate_freq=2, jit=off
opcache runtime hit rate: UNKNOWN (no status scrape of /fpm-status from remote; localhost-only)
```
Workers sampled: **6** skinbase processes, **avg RSS ~95 MB**, total ~570 MB.
Theoretical ceiling: `14 × 256 MB = 3.6 GB` if every worker hit `memory_limit`. Typical: `14 × 95 MB ≈ 1.3 GB`.
Available RAM 8.6 GiB **but swap already full** — do not raise max_children in M0.
---
## 13–15. Laravel Production Configuration (PRODUCTION SERVER)
`php artisan about` + tinker `config()` (no `.env` dump):
```text
APP_ENV=production
APP_DEBUG=false
CACHE_STORE=redis
SESSION_DRIVER=redis
QUEUE_CONNECTION=redis
SCOUT_DRIVER=meilisearch
FILESYSTEM_DISK=local
UPLOAD_QUEUE_DERIVATIVES=false
DOWNLOAD_ACCEL_ENABLED=false
DOWNLOAD_ACCEL_PATH=/internal/originals (path set, flag off, nginx location missing)
SITEMAPS_BUILD_ON_REQUEST=true
SITEMAPS_FALLBACK_TO_LIVE_BUILD=true
SITEMAPS_PREGENERATED_ENABLED=true
SITEMAPS_PREFER_PUBLISHED_RELEASE=true
SITEMAPS_STATIC_PUBLISH_ENABLED=true
REDIS_CLIENT=predis
homepage.cache_store=homepage
homepage.guest_payload_ttl_seconds=1800
recommendations.queue=default
vision.queue=default
discovery.queue=default
broadcasting=reverb
horizon.path=horizon
Sentry enabled, sample rate errors 100%, performance 100%
```
Laravel caches: **config, events, routes, views ALL CACHED**.
Debugbar/Telescope/Clockwork routes: **none**. Composer `packages-dev=0` → **--no-dev install**.
---
## 16–22. MySQL (PRODUCTION SERVER)
```text
location: local mysqld
version: 8.4.11-11
database: skinbase
size: 2758.2 MB (data 1230.4 + indexes 1527.8)
buffer pool: 10 GB, 10 instances
max_conn: 120 (Max_used 25)
slow_query_log: ON, long_query_time=0.5s, file /var/log/mysql/slow.log (not readable here)
Slow_queries: 11342 in Uptime 97330s (~27h) ≈ 0.12/s
Threads: connected 3, running 2
InnoDB hit: 1 - 103791/18420885623 ≈ 99.999%
Rows read: 9.70e9 in ~27h (high — job/scan load)
tmp tables: 455567 created, 21 on disk
```
Largest tables (information_schema estimates):
| Table | ~rows | total MB |
| ----- | -----: | -------: |
| artwork_metric_snapshots_hourly | 8.62M | **1855.5** |
| artwork_downloads | 579k | 167.3 |
| forum_bot_logs | 267k | 141.8 |
| artworks | 50k | 94.2 |
| user_activities | 240k | 69.3 |
| artwork_comments | 172k | 62.2 |
| artwork_view_events | 161k | 25.4 |
| rank_artwork_scores | 50k | 22.4 |
| rec_item_pairs | 74k | 16.0 |
| artwork_stats | 49k | 15.2 |
`cache`/`sessions` tables nearly empty (drivers are Redis). `jobs` table empty (Redis queues).
### Indexes (batch1) — PRODUCTION
| Index | Status |
| ----- | ------ |
| artworks.idx_public_approved_published_id | **PRESENT** |
| artworks.idx_public_approved_user_id | **PRESENT** |
| artworks FULLTEXT title+description | **PRESENT** |
| snapshots.idx_bucket_artwork | **PRESENT** |
| rank_artwork_scores idx_mv_trending/new_hot/best | **PRESENT** |
| tags.artworks_count + idx_tags_artworks_count | **PRESENT** |
### DB-001 re-evaluation
```text
Classification: UNKNOWN as “still 78%”, NOT “unindexed public scans”
```
April indexes **are deployed**. Dominant storage is hourly snapshots ≈ 50k artworks × 24h × 7d. Heat/ranking jobs that join this table can still dominate `Innodb_rows_read` even with `idx_bucket_artwork`. Cannot confirm fingerprints without `slow.log` / performance_schema (denied).
Do not add more indexes in M0.
---
## 23–25. Redis / Cache / Sessions (PRODUCTION SERVER)
```text
redis_version: 8.10.1
used: 877.55 MB (peak 917 MB)
maxmemory: 2.00 G, policy allkeys-lru
evicted_keys: 0
clients: 20
ops/sec: 44 (instant)
keyspace hits/misses: 1.40M / 1.30M (~52% hit)
db0: 1932 keys; db1: 6641; db2: 55176 (all with TTL)
```
| System | Redis in production? |
| ------ | -------------------- |
| Cache default | **YES** |
| Homepage store `homepage` | configured (failover redis→database) |
| Sessions | **YES** |
| Queues / Horizon | **YES** |
| Presence | code uses Redis (not separately counted) |
| Stats deltas | flush command scheduled; view path still MySQL sync |
CACHE-001 from the code audit (**database cache default**) is **NOT ACTIVE in production**.
---
## 26–32. Horizon / Queues / QUEUE-001 (PRODUCTION SERVER)
Horizon **running**. Supervisors:
| Supervisor | queues | timeout | maxProcesses |
| ---------- | ------ | ------: | -----------: |
| supervisor-default | search, default | **960** | 5 |
| supervisor-messaging | broadcasts, notifications | **90** | 3 |
`recommendations.queue` / `vision.queue` / `discovery.queue` = **`default`** → consumed by 960s workers.
Standalone supervisor `queue:work --timeout=90` from `deploy/supervisor/skinbase-queue.conf` is **NOT the production worker**.
Failed jobs: **329** rows (`queue:failed` listed many; table_rows estimate 159). Dominant:
```text
RecComputeSimilarByBehaviorJob ~02:18 daily MaxAttemptsExceededException
RecComputeSimilarHybridJob ~02:31 daily MaxAttemptsExceededException
```
Horizon worker `tries=1`, job `$timeout=900`, worker timeout 960. Failures are **not** explained by the old 90s worker. Likely 900s job timeout, 128 MB worker memory, or lock/overlap. **Do not retry in M0.**
Queue LLEN (Redis): all listed queues 0 except **`queues:mail` = 3**. Horizon does **not** listen to `mail`. Mail jobs can stall.
```text
QUEUE-001 (90 vs 900): MITIGATED
Nightly rec job failures: ACTIVE P0 (related, different mechanism)
mail queue unconsumed: P1
```
---
## 31–32. Scheduler (PRODUCTION SERVER)
`php artisan schedule:list` matches `routes/console.php` (generate sitemaps 10:30/22:30, rec jobs 02:00–02:30, etc.).
**How `schedule:run` is invoked:** user crontab empty; sudo crontab **not readable**; no systemd timer named laravel/schedule. **Empirically running** (sitemap shards 10:30 today; rec jobs fail 02:18/02:31 daily). Likely root crontab. Do not change.
02:00–05:00 window (code + evidence it fires): rec jobs, ranking/heat, analytics, prune snapshots, sitemap validate — plus nightly rec **failures**.
---
## 33. Meilisearch (PRODUCTION SERVER)
```text
process: /usr/local/bin/meilisearch --config-file-path /etc/meilisearch.toml
health: {"status":"available"}
version/stats: 401 without key (not printed)
RSS: ~401 MB
```
No reindex.
---
## 34–36. Storage / Downloads / Sitemaps
- `FILESYSTEM_DISK=local`; CDN for public derivatives.
- Download accel **off**; originals would stream via PHP if used.
- Sitemaps live under shared `.../shared/public/sitemaps/` (20M, 27 xml files).
- **Shards + academy + users + forum-threads mtime 2026-08-23 10:30–10:31**
- **`sitemap.xml` mtime 2026-05-13 21:39, 1906 bytes** — generate/publish does **not** update the index file.
- May index omits academy-* families that now exist on disk.
SITEMAP-001 live-build-on-missing: files exist, so crawlers are not building XML in PHP for the index. `@php` named location missing anyway.
```text
SITEMAP-001 original (PHP live-build stampede): NOT ACTIVE right now
Stale sitemap index vs daily shards: ACTIVE (SEO/ops)
```
---
## 37–38. Views / Downloads
```text
artwork_view_events last 1h: 72
artwork_view_events last 24h: 2666
artwork_downloads last 24h: 4718
view_events table: ~161k rows, 25.4 MB
downloads table: ~579k rows, 167 MB
```
STAT-001 architecture (sync INSERT+UPDATE, `defer=false`) is **still the code on this release**. At **~0.03 views/s** it is not the current DB bottleneck. Classify **P1** (architecture / growth), not active P0 load.
Downloads were not fetched (would write).
---
## 39–47. Latency / Traffic / Resources
Origin access-log QPS not aggregated (would include IPs — skipped). Edge n=1 TTFBs from earlier remain **PUBLIC** evidence, not origin APM.
FPM slowlog: 139k historical lines, last file mtime Aug 23 05:05, timeout 10s. Scripts are `index.php` only (no URI dump here).
Resource baseline: see §4. **Swap full = CONCERNING.**
---
## 48. Lighthouse / CWV
Not re-run. Prior lab file is skinbase.top 2026-03-23 — do not use as this baseline.
---
## 49. Observability
- Sentry **on**, 100% error **and** performance sample (P1 cost).
- Netdata on host.
- Horizon dashboard path `horizon` (auth not verified; do not probe).
- `/fpm-status` localhost only.
- `/stats/` basic-auth (not accessed).
- MySQL slow.log exists but unreadable to this SSH user.
---
## 50. Re-evaluated P0 Findings
| ID | Verdict | Why |
| -- | ------- | --- |
| QUEUE-001 90s worker vs 900s job | **MITIGATED** | Horizon default timeout **960**; rec queues aliased to `default` |
| RecCompute* nightly MaxAttemptsExceeded | **ACTIVE P0** | every night 02:18/02:31 on `redis@default`; 329 failed jobs |
| STAT-001 sync view writes | **P1** at current 2666/day | still sync in config/code; not the load leader |
| SITEMAP-001 request-time build | **NOT ACTIVE** | static files present; `@php` missing |
| Stale sitemap.xml (May vs Aug shards) | **ACTIVE P1** (SEO) | generate writes children, not index |
| DB-001 unindexed aggregates | **MITIGATED for missing indexes** | batch1 **PRESENT**. Residual scan cost **UNKNOWN**; snapshots 1.85 GB |
---
## 51. Re-evaluated P1 Findings
- Production **Redis** cache/session (repo default database is wrong for prod).
- Upload 20M/24M FPM vs app 50M/200M.
- `UPLOAD_QUEUE_DERIVATIVES=false` — sync image work on publish.
- `mail` Redis queue not consumed by Horizon.
- Sentry performance sampling 100%.
- CF HTML DYNAMIC despite public Cache-Control.
- Asset TTL 30d not 1y immutable.
- Swap 2G full; shared ML/Docker host.
- FPM slowlog 10s historically large.
- Local git **newer** than production release (f52879ed vs 5af95f65).
- Session cookies on `/search` and 404 (code-consistent).
- Duplicate security headers (app middleware + nginx snippet).
---
## 52. Baseline Metric Table
| Metric | Value | Source |
| ------ | ----- | ------ |
| DB size | 2758 MB | information_schema |
| Snapshots table | 1855 MB / ~8.6M rows | information_schema |
| Artworks | ~50k / 94 MB | information_schema |
| Views 24h | 2666 | count on 161k table |
| Downloads 24h | 4718 | count |
| Slow_queries / 27h | 11342 | SHOW STATUS |
| InnoDB pool hit | ~100% | STATUS |
| Redis used | 878 MB / 2 GB | INFO |
| Redis hit ratio | ~52% | INFO |
| FPM workers | 6 / max 14, ~95 MB RSS | ps |
| Horizon | running, timeout 960/90 | ps + artisan |
| Failed jobs | 329 | DB |
| Homepage TTFB (edge) | 146 ms n=1 | PUBLIC |
| Production release | 20260801-172803-5af95f65-dirty | filesystem |
---
## 53. Production Risk Matrix
| Risk | Evidence | Impact | Confidence |
| ---- | -------- | ------ | ---------- |
| Nightly rec jobs never succeed | failed_jobs every day | stale similar-art | high |
| Hourly snapshot table 1.85 GB | information_schema | heat/rank I/O | high |
| Swap full + co-tenant ML | free/ps | latency spikes | high |
| Stale sitemap index | mtime May 13 vs shards today | crawl/SEO | high |
| mail queue not consumed | LLEN=3, Horizon queues | delayed mail | high |
| Missing @php location | nginx vhost | missing shard → error not Laravel | high |
| Cannot read slow.log | permissions | DB-001 residual | high |
---
## 54. Optimization Targets (do not implement now)
1. Rec job failures (timeout/memory/overlap) — after measuring why MaxAttemptsExceeded.
2. Snapshot retention / heat query EXPLAIN once slow.log is readable.
3. Publish `sitemap.xml` index with shards.
4. Consume `mail` or stop using that queue name.
5. Download X-Accel if originals still hit FPM.
6. Align upload limits.
7. Memory/swap / co-tenancy.
8. Sentry sample rates.
9. CF HTML cache policy (product decision).
---
## 55. M1–M4 Readiness
```text
M1 planning: READY (production drivers known; P0 recast)
M1 implementation: not this milestone
Blockers for DB-001 closeout: mysql slow.log or GRANTs on performance_schema
```
---
## 56. Unknowns
- Root crontab contents (scheduler invocation path)
- MySQL slow.log fingerprints **now**
- OPcache hit rate (FPM status localhost-only)
- Meilisearch document counts (auth)
- Peak QPS / access-log (IPs omitted on purpose)
- Why rec jobs MaxAttemptsExceeded (timeout vs memory vs exception) — do not dump payloads
- HTTP/3 to origin (not configured)
- iostat (command missing)
---
## 57. Repository Changes
```text
Created earlier:
- scripts/collect-production-baseline.sh
Updated:
- docs/skinbase-production-baseline.md
Modified application:
- none
Database changes:
- none
Server configuration changes:
- none
Service restarts:
- none
Cache clears:
- none
Queue operations:
- none (failed jobs listed only, not retried)
```