# Skinbase.org — M0 Production Baseline & Runtime Audit ```text STATUS: COMPLETE ``` Primary origin runtime unknowns were resolved via **read-only `ssh server3`**. Some items remain UNKNOWN (MySQL `slow.log` not readable by this user; root crontab not readable without sudo; `performance_schema` denied to `skinbase@localhost`; `iostat` not installed). **Nothing was modified on server3.** No restarts, no cache clears, no SQL writes, no deploys. Evidence classes used below: ```text LOCAL CODE EVIDENCE PRODUCTION SERVER EVIDENCE PUBLIC CLOUDFLARE/HTTP EVIDENCE ``` --- ## Executive Summary ```text Production OS: Debian 13 (trixie), kernel 6.12.96+deb13-cloud-amd64, x86_64 CPU: AMD EPYC, 8 cores / 8 threads (1 socket) RAM: 23 GiB; ~8.6 GiB available; SWAP 2.0 GiB **almost fully used** Storage: 400G virtio disk `/dev/sda1` ext4 ~185G/394G (49%); rotational=0 (SSD/NVMe-like) nginx: 1.26.3; HTTP/2 on; gzip+brotli in nginx.conf; vhost skinbase.org.conf PHP: 8.4.24 (FPM + CLI) PHP-FPM: pool `skinbase`, dynamic, max_children=14, memory_limit=256M, terminate=120s MySQL: 8.4.11-11 local; DB `skinbase` 2758 MB; buffer pool 10 GB Redis: 8.10.1 local 127.0.0.1:6379; 878 MB used / 2 GB max; allkeys-lru Meilisearch: local systemd; health available; RSS ~401 MB Queue manager: Supervisor → `php artisan horizon` (NOT standalone queue:work) Cache store: **redis** (PRODUCTION; not repo default database) Session store: **redis** CDN: Cloudflare + cdn.skinbase.org RGW/S3 Download acceleration: **OFF** (`download_accel_enabled=false`; no X-Accel location in vhost) Sitemap static serving: **ACTIVE** (nginx try_files + shared files). Index file stale (May 13). Debug mode: OFF Debugbar production: NOT PRESENT (packages-dev=0, no debugbar routes) ``` ```text PRODUCTION_APP_PATH=/opt/www/virtual/SkinbaseNova → symlink → /opt/www/virtual/SkinbaseNova.releases/current → /opt/www/virtual/SkinbaseNova.releases/releases/20260801-172803-5af95f65-dirty Local commit: f52879edbb19ecc5d3e807bd4872016f3071dad7 (develop) Production commit: 5af95f65 (release name 5af95f65-dirty, 2026-08-01) Same revision: no Production git: not a git checkout (no .git in release) ``` ```text Active P0 findings: QUEUE-001 (90s vs 900s): MITIGATED — Horizon default workers timeout=960 QUEUE rec-job failures: ACTIVE P0 — RecComputeSimilar* fail every night MaxAttemptsExceeded SITEMAP-001 live-build: NOT the current incident — static files exist SITEMAP index stale: ACTIVE P1/P0-SEO — sitemap.xml mtime 2026-05-13 while shards refresh daily STAT-001: NOT currently P0 at 2666 views/24h; design still sync writes (P1) DB-001: indexes PRESENT; residual UNKNOWN without slow.log. Snapshot table 1.85 GB is the dominant store. Active P1 findings: swap exhaustion; mail queue unconsumed (LLEN=3); upload 20M vs app 50M; Sentry traces 100%; FPM 10s slowlog volume; CF HTML DYNAMIC; asset TTL 30d not 1y immutable Major current bottleneck: MySQL row-read volume + 1.85 GB hourly snapshots; nightly rec jobs failing; memory pressure (swap full) Largest remaining unknown: contents of /var/log/mysql/slow.log (permission) M1 readiness: YES for planning. Do not implement in M0. ``` --- ## 1. Repository State (LOCAL) ```text branch: develop HEAD: f52879edbb19ecc5d3e807bd4872016f3071dad7 commit: f52879ed Current state with latest updates tree: clean at start of this SSH pass (ahead of origin/develop by 1) ``` No reset/stash/clean on local or production. --- ## 2. Production Host (PRODUCTION SERVER) ```text hostname: server3 FQDN: server3.klevze.si OS: Debian GNU/Linux 13 (trixie) 13.6 kernel: 6.12.96+deb13-cloud-amd64 arch: x86_64 uptime: 24 days, 16:58 (sampled 2026-08-23 13:43 CEST) ``` --- ## 3. Hardware Baseline (PRODUCTION SERVER) ```text CPU model: AMD EPYC Processor (with IBPB) sockets: 1 cores: 8 threads: 8 (1 thread/core) RAM: 23 GiB swap: 2.0 GiB, ~2.0 GiB used, 2.7 MiB free disk: /dev/sda 400G → sda1 399.9G on / filesystem: ~394G, 185G used, 49% rotational: 0 → SSD/cloud SSD (not HDD) ``` `dmidecode` not used (would need sudo). --- ## 4. Current System Load (PRODUCTION SERVER) Short snapshot only (limitation: not peak-hour). ```text load average: 0.85, 2.47, 2.88 vmstat 5s: idle ~76–83%, wa ~0–1%, r=3–6 CPU notables: mysqld ~20%, containerd/dockerd ~45% each (other workloads on same host), redis ~5% RAM: 14 GiB used + 7 GiB cache; 8.6 GiB available swap: fully used ``` Largest RSS: | Process | RSS | Note | | ------- | --: | ---- | | mysqld | ~7.7 GiB | 32.7% | | python uvicorn × several | up to 1.3 GiB | vision/ML stack | | clamd | ~1.1 GiB | | | meilisearch | ~397–401 MiB | | | redis | ~164 MiB | | | php-fpm skinbase | ~95–100 MiB each | 6 workers sampled | | node ssr.js | ~122 MiB | | Classification: **WATCH / CONCERNING** (swap exhausted; shared host with Docker/ML/Qdrant/Gitea/Crowdsec). Not a Skinbase-only box. --- ## 5. Process Inventory (PRODUCTION SERVER) | Component | Reality | | --------- | ------- | | Horizon | **YES** — Supervisor `skinbase-horizon`, pid running ~1d | | standalone `queue:work` | **NO** | | Supervisor | **YES** — horizon + ssr | | systemd queue workers | **NO** | | Inertia SSR | **YES** — `node bootstrap/ssr/ssr.js` as www-data | | Reverb | **YES** — systemd `skinbase-reverb.service`, 127.0.0.1:8080 | | Meilisearch | **YES** — local systemd | | Redis | **local** 127.0.0.1:6379 | | MySQL | **local** `/usr/sbin/mysqld` | | PHP-FPM | php8.4, pool `skinbase` + unused pool `www` | | nginx | master + 8 workers, www-data | | Docker/ML | uvicorn:8000, qdrant, gitea — **co-tenant** | --- ## 6. nginx (PRODUCTION SERVER) ```text version: nginx/1.26.3 vhost: /etc/nginx/sites-enabled/skinbase.org.conf root: /opt/www/virtual/SkinbaseNova/public listen: 80 → 301 HTTPS; 443 ssl; http2 on PHP: unix:/run/php/php8.4-fpm-skinbase.sock client_max_body_size: 24m gzip: on (nginx.conf) brotli: on (nginx.conf) real IP: /etc/nginx/conf.d/00-cloudflare-realip.conf ACTIVE search: location = /search + conf.d/13-skinbase-search-protection.conf (20r/m, abusive query map) HTTP/3: not in this vhost (Cloudflare alt-svc h3 is edge-only) ``` Snippet vs live: | Snippet | Status | | ------- | ------ | | sitemaps.conf (`max-age=21600`, try_files) | **ACTIVE** (inlined in vhost) | | static-cache.conf 1y immutable `/build/assets` | **NOT ACTIVE** — generic `expires 30d` for css/js/images | | download-accel.conf | **NOT ACTIVE** | | search-rate-limit.conf (repo) | **EQUIVALENT ACTIVE** as `13-skinbase-search-protection.conf` | | upstream-error-pages.conf | **NOT seen** in vhost | **Bug/risk:** sitemap locations use `try_files $uri @php` but **no `location @php`** exists in the vhost. Missing files will not fall through to Laravel as the snippet comments claim. No `nginx -T` via sudo (password required). Readable vhost was enough. --- ## 7–9. HTTP / CDN (PUBLIC CLOUDFLARE/HTTP + PRODUCTION) Unchanged from edge probes, now correlated with origin: - Homepage Cache-Control matches Laravel guest headers; CF **DYNAMIC** - Build assets CF **HIT**, origin `expires 30d` - CDN webp `max-age=31536000` HIT + Polish - Brotli **CONFIRMED** at edge; origin brotli **on** - HTTP/2 **CONFIRMED** at origin (`http2 on`); this Windows curl was HTTP/1.1 to CF - HTTP/3 **offered by CF** (`alt-svc`), not configured on origin vhost --- ## 10–12. PHP / OPcache / FPM (PRODUCTION SERVER) ```text PHP: 8.4.24 FPM master: /etc/php/8.4/fpm/php-fpm.conf pool: [skinbase] user=skinbase listen: /run/php/php8.4-fpm-skinbase.sock pm: dynamic max_children: 14 start_servers: 4 min/max spare: 2 / 6 max_requests: 500 request_terminate_timeout: 120s slowlog: 10s → /var/log/php8.4-fpm-skinbase-slow.log (23 MB, 139556 lines) memory_limit: 256M (pool) max_execution: 60s (pool) upload/post: 20M / 24M (pool) ← app allows 50M images / 200M archives CLI upload_max: 2M (irrelevant to FPM) extensions: opcache, redis, pcntl, posix, intl, mbstring, curl, gd, pdo_mysql imagick: NOT in php -m opcache (FPM php.ini): memory 256MB, interned 16, max_files 20000, validate_timestamps=1, revalidate_freq=2, jit=off opcache runtime hit rate: UNKNOWN (no status scrape of /fpm-status from remote; localhost-only) ``` Workers sampled: **6** skinbase processes, **avg RSS ~95 MB**, total ~570 MB. Theoretical ceiling: `14 × 256 MB = 3.6 GB` if every worker hit `memory_limit`. Typical: `14 × 95 MB ≈ 1.3 GB`. Available RAM 8.6 GiB **but swap already full** — do not raise max_children in M0. --- ## 13–15. Laravel Production Configuration (PRODUCTION SERVER) `php artisan about` + tinker `config()` (no `.env` dump): ```text APP_ENV=production APP_DEBUG=false CACHE_STORE=redis SESSION_DRIVER=redis QUEUE_CONNECTION=redis SCOUT_DRIVER=meilisearch FILESYSTEM_DISK=local UPLOAD_QUEUE_DERIVATIVES=false DOWNLOAD_ACCEL_ENABLED=false DOWNLOAD_ACCEL_PATH=/internal/originals (path set, flag off, nginx location missing) SITEMAPS_BUILD_ON_REQUEST=true SITEMAPS_FALLBACK_TO_LIVE_BUILD=true SITEMAPS_PREGENERATED_ENABLED=true SITEMAPS_PREFER_PUBLISHED_RELEASE=true SITEMAPS_STATIC_PUBLISH_ENABLED=true REDIS_CLIENT=predis homepage.cache_store=homepage homepage.guest_payload_ttl_seconds=1800 recommendations.queue=default vision.queue=default discovery.queue=default broadcasting=reverb horizon.path=horizon Sentry enabled, sample rate errors 100%, performance 100% ``` Laravel caches: **config, events, routes, views ALL CACHED**. Debugbar/Telescope/Clockwork routes: **none**. Composer `packages-dev=0` → **--no-dev install**. --- ## 16–22. MySQL (PRODUCTION SERVER) ```text location: local mysqld version: 8.4.11-11 database: skinbase size: 2758.2 MB (data 1230.4 + indexes 1527.8) buffer pool: 10 GB, 10 instances max_conn: 120 (Max_used 25) slow_query_log: ON, long_query_time=0.5s, file /var/log/mysql/slow.log (not readable here) Slow_queries: 11342 in Uptime 97330s (~27h) ≈ 0.12/s Threads: connected 3, running 2 InnoDB hit: 1 - 103791/18420885623 ≈ 99.999% Rows read: 9.70e9 in ~27h (high — job/scan load) tmp tables: 455567 created, 21 on disk ``` Largest tables (information_schema estimates): | Table | ~rows | total MB | | ----- | -----: | -------: | | artwork_metric_snapshots_hourly | 8.62M | **1855.5** | | artwork_downloads | 579k | 167.3 | | forum_bot_logs | 267k | 141.8 | | artworks | 50k | 94.2 | | user_activities | 240k | 69.3 | | artwork_comments | 172k | 62.2 | | artwork_view_events | 161k | 25.4 | | rank_artwork_scores | 50k | 22.4 | | rec_item_pairs | 74k | 16.0 | | artwork_stats | 49k | 15.2 | `cache`/`sessions` tables nearly empty (drivers are Redis). `jobs` table empty (Redis queues). ### Indexes (batch1) — PRODUCTION | Index | Status | | ----- | ------ | | artworks.idx_public_approved_published_id | **PRESENT** | | artworks.idx_public_approved_user_id | **PRESENT** | | artworks FULLTEXT title+description | **PRESENT** | | snapshots.idx_bucket_artwork | **PRESENT** | | rank_artwork_scores idx_mv_trending/new_hot/best | **PRESENT** | | tags.artworks_count + idx_tags_artworks_count | **PRESENT** | ### DB-001 re-evaluation ```text Classification: UNKNOWN as “still 78%”, NOT “unindexed public scans” ``` April indexes **are deployed**. Dominant storage is hourly snapshots ≈ 50k artworks × 24h × 7d. Heat/ranking jobs that join this table can still dominate `Innodb_rows_read` even with `idx_bucket_artwork`. Cannot confirm fingerprints without `slow.log` / performance_schema (denied). Do not add more indexes in M0. --- ## 23–25. Redis / Cache / Sessions (PRODUCTION SERVER) ```text redis_version: 8.10.1 used: 877.55 MB (peak 917 MB) maxmemory: 2.00 G, policy allkeys-lru evicted_keys: 0 clients: 20 ops/sec: 44 (instant) keyspace hits/misses: 1.40M / 1.30M (~52% hit) db0: 1932 keys; db1: 6641; db2: 55176 (all with TTL) ``` | System | Redis in production? | | ------ | -------------------- | | Cache default | **YES** | | Homepage store `homepage` | configured (failover redis→database) | | Sessions | **YES** | | Queues / Horizon | **YES** | | Presence | code uses Redis (not separately counted) | | Stats deltas | flush command scheduled; view path still MySQL sync | CACHE-001 from the code audit (**database cache default**) is **NOT ACTIVE in production**. --- ## 26–32. Horizon / Queues / QUEUE-001 (PRODUCTION SERVER) Horizon **running**. Supervisors: | Supervisor | queues | timeout | maxProcesses | | ---------- | ------ | ------: | -----------: | | supervisor-default | search, default | **960** | 5 | | supervisor-messaging | broadcasts, notifications | **90** | 3 | `recommendations.queue` / `vision.queue` / `discovery.queue` = **`default`** → consumed by 960s workers. Standalone supervisor `queue:work --timeout=90` from `deploy/supervisor/skinbase-queue.conf` is **NOT the production worker**. Failed jobs: **329** rows (`queue:failed` listed many; table_rows estimate 159). Dominant: ```text RecComputeSimilarByBehaviorJob ~02:18 daily MaxAttemptsExceededException RecComputeSimilarHybridJob ~02:31 daily MaxAttemptsExceededException ``` Horizon worker `tries=1`, job `$timeout=900`, worker timeout 960. Failures are **not** explained by the old 90s worker. Likely 900s job timeout, 128 MB worker memory, or lock/overlap. **Do not retry in M0.** Queue LLEN (Redis): all listed queues 0 except **`queues:mail` = 3**. Horizon does **not** listen to `mail`. Mail jobs can stall. ```text QUEUE-001 (90 vs 900): MITIGATED Nightly rec job failures: ACTIVE P0 (related, different mechanism) mail queue unconsumed: P1 ``` --- ## 31–32. Scheduler (PRODUCTION SERVER) `php artisan schedule:list` matches `routes/console.php` (generate sitemaps 10:30/22:30, rec jobs 02:00–02:30, etc.). **How `schedule:run` is invoked:** user crontab empty; sudo crontab **not readable**; no systemd timer named laravel/schedule. **Empirically running** (sitemap shards 10:30 today; rec jobs fail 02:18/02:31 daily). Likely root crontab. Do not change. 02:00–05:00 window (code + evidence it fires): rec jobs, ranking/heat, analytics, prune snapshots, sitemap validate — plus nightly rec **failures**. --- ## 33. Meilisearch (PRODUCTION SERVER) ```text process: /usr/local/bin/meilisearch --config-file-path /etc/meilisearch.toml health: {"status":"available"} version/stats: 401 without key (not printed) RSS: ~401 MB ``` No reindex. --- ## 34–36. Storage / Downloads / Sitemaps - `FILESYSTEM_DISK=local`; CDN for public derivatives. - Download accel **off**; originals would stream via PHP if used. - Sitemaps live under shared `.../shared/public/sitemaps/` (20M, 27 xml files). - **Shards + academy + users + forum-threads mtime 2026-08-23 10:30–10:31** - **`sitemap.xml` mtime 2026-05-13 21:39, 1906 bytes** — generate/publish does **not** update the index file. - May index omits academy-* families that now exist on disk. SITEMAP-001 live-build-on-missing: files exist, so crawlers are not building XML in PHP for the index. `@php` named location missing anyway. ```text SITEMAP-001 original (PHP live-build stampede): NOT ACTIVE right now Stale sitemap index vs daily shards: ACTIVE (SEO/ops) ``` --- ## 37–38. Views / Downloads ```text artwork_view_events last 1h: 72 artwork_view_events last 24h: 2666 artwork_downloads last 24h: 4718 view_events table: ~161k rows, 25.4 MB downloads table: ~579k rows, 167 MB ``` STAT-001 architecture (sync INSERT+UPDATE, `defer=false`) is **still the code on this release**. At **~0.03 views/s** it is not the current DB bottleneck. Classify **P1** (architecture / growth), not active P0 load. Downloads were not fetched (would write). --- ## 39–47. Latency / Traffic / Resources Origin access-log QPS not aggregated (would include IPs — skipped). Edge n=1 TTFBs from earlier remain **PUBLIC** evidence, not origin APM. FPM slowlog: 139k historical lines, last file mtime Aug 23 05:05, timeout 10s. Scripts are `index.php` only (no URI dump here). Resource baseline: see §4. **Swap full = CONCERNING.** --- ## 48. Lighthouse / CWV Not re-run. Prior lab file is skinbase.top 2026-03-23 — do not use as this baseline. --- ## 49. Observability - Sentry **on**, 100% error **and** performance sample (P1 cost). - Netdata on host. - Horizon dashboard path `horizon` (auth not verified; do not probe). - `/fpm-status` localhost only. - `/stats/` basic-auth (not accessed). - MySQL slow.log exists but unreadable to this SSH user. --- ## 50. Re-evaluated P0 Findings | ID | Verdict | Why | | -- | ------- | --- | | QUEUE-001 90s worker vs 900s job | **MITIGATED** | Horizon default timeout **960**; rec queues aliased to `default` | | RecCompute* nightly MaxAttemptsExceeded | **ACTIVE P0** | every night 02:18/02:31 on `redis@default`; 329 failed jobs | | STAT-001 sync view writes | **P1** at current 2666/day | still sync in config/code; not the load leader | | SITEMAP-001 request-time build | **NOT ACTIVE** | static files present; `@php` missing | | Stale sitemap.xml (May vs Aug shards) | **ACTIVE P1** (SEO) | generate writes children, not index | | DB-001 unindexed aggregates | **MITIGATED for missing indexes** | batch1 **PRESENT**. Residual scan cost **UNKNOWN**; snapshots 1.85 GB | --- ## 51. Re-evaluated P1 Findings - Production **Redis** cache/session (repo default database is wrong for prod). - Upload 20M/24M FPM vs app 50M/200M. - `UPLOAD_QUEUE_DERIVATIVES=false` — sync image work on publish. - `mail` Redis queue not consumed by Horizon. - Sentry performance sampling 100%. - CF HTML DYNAMIC despite public Cache-Control. - Asset TTL 30d not 1y immutable. - Swap 2G full; shared ML/Docker host. - FPM slowlog 10s historically large. - Local git **newer** than production release (f52879ed vs 5af95f65). - Session cookies on `/search` and 404 (code-consistent). - Duplicate security headers (app middleware + nginx snippet). --- ## 52. Baseline Metric Table | Metric | Value | Source | | ------ | ----- | ------ | | DB size | 2758 MB | information_schema | | Snapshots table | 1855 MB / ~8.6M rows | information_schema | | Artworks | ~50k / 94 MB | information_schema | | Views 24h | 2666 | count on 161k table | | Downloads 24h | 4718 | count | | Slow_queries / 27h | 11342 | SHOW STATUS | | InnoDB pool hit | ~100% | STATUS | | Redis used | 878 MB / 2 GB | INFO | | Redis hit ratio | ~52% | INFO | | FPM workers | 6 / max 14, ~95 MB RSS | ps | | Horizon | running, timeout 960/90 | ps + artisan | | Failed jobs | 329 | DB | | Homepage TTFB (edge) | 146 ms n=1 | PUBLIC | | Production release | 20260801-172803-5af95f65-dirty | filesystem | --- ## 53. Production Risk Matrix | Risk | Evidence | Impact | Confidence | | ---- | -------- | ------ | ---------- | | Nightly rec jobs never succeed | failed_jobs every day | stale similar-art | high | | Hourly snapshot table 1.85 GB | information_schema | heat/rank I/O | high | | Swap full + co-tenant ML | free/ps | latency spikes | high | | Stale sitemap index | mtime May 13 vs shards today | crawl/SEO | high | | mail queue not consumed | LLEN=3, Horizon queues | delayed mail | high | | Missing @php location | nginx vhost | missing shard → error not Laravel | high | | Cannot read slow.log | permissions | DB-001 residual | high | --- ## 54. Optimization Targets (do not implement now) 1. Rec job failures (timeout/memory/overlap) — after measuring why MaxAttemptsExceeded. 2. Snapshot retention / heat query EXPLAIN once slow.log is readable. 3. Publish `sitemap.xml` index with shards. 4. Consume `mail` or stop using that queue name. 5. Download X-Accel if originals still hit FPM. 6. Align upload limits. 7. Memory/swap / co-tenancy. 8. Sentry sample rates. 9. CF HTML cache policy (product decision). --- ## 55. M1–M4 Readiness ```text M1 planning: READY (production drivers known; P0 recast) M1 implementation: not this milestone Blockers for DB-001 closeout: mysql slow.log or GRANTs on performance_schema ``` --- ## 56. Unknowns - Root crontab contents (scheduler invocation path) - MySQL slow.log fingerprints **now** - OPcache hit rate (FPM status localhost-only) - Meilisearch document counts (auth) - Peak QPS / access-log (IPs omitted on purpose) - Why rec jobs MaxAttemptsExceeded (timeout vs memory vs exception) — do not dump payloads - HTTP/3 to origin (not configured) - iostat (command missing) --- ## 57. Repository Changes ```text Created earlier: - scripts/collect-production-baseline.sh Updated: - docs/skinbase-production-baseline.md Modified application: - none Database changes: - none Server configuration changes: - none Service restarts: - none Cache clears: - none Queue operations: - none (failed jobs listed only, not retried) ```