Keep similar-ai from tripping the global circuit on a lone URL 502, clamp Qdrant search to 100, and add Server-Timing plus slow-request logging. Studio shared props, Academy S3 exists caching, heat chunking, and Redis/scheduler hygiene stay in this rollout.
21 KiB
Skinbase.org — M0 Production Baseline & Runtime Audit
STATUS: COMPLETE
Primary origin runtime unknowns were resolved via read-only ssh server3. Some items remain UNKNOWN (MySQL slow.log not readable by this user; root crontab not readable without sudo; performance_schema denied to skinbase@localhost; iostat not installed).
Nothing was modified on server3. No restarts, no cache clears, no SQL writes, no deploys.
Evidence classes used below:
LOCAL CODE EVIDENCE
PRODUCTION SERVER EVIDENCE
PUBLIC CLOUDFLARE/HTTP EVIDENCE
Executive Summary
Production OS: Debian 13 (trixie), kernel 6.12.96+deb13-cloud-amd64, x86_64
CPU: AMD EPYC, 8 cores / 8 threads (1 socket)
RAM: 23 GiB; ~8.6 GiB available; SWAP 2.0 GiB **almost fully used**
Storage: 400G virtio disk `/dev/sda1` ext4 ~185G/394G (49%); rotational=0 (SSD/NVMe-like)
nginx: 1.26.3; HTTP/2 on; gzip+brotli in nginx.conf; vhost skinbase.org.conf
PHP: 8.4.24 (FPM + CLI)
PHP-FPM: pool `skinbase`, dynamic, max_children=14, memory_limit=256M, terminate=120s
MySQL: 8.4.11-11 local; DB `skinbase` 2758 MB; buffer pool 10 GB
Redis: 8.10.1 local 127.0.0.1:6379; 878 MB used / 2 GB max; allkeys-lru
Meilisearch: local systemd; health available; RSS ~401 MB
Queue manager: Supervisor → `php artisan horizon` (NOT standalone queue:work)
Cache store: **redis** (PRODUCTION; not repo default database)
Session store: **redis**
CDN: Cloudflare + cdn.skinbase.org RGW/S3
Download acceleration: **OFF** (`download_accel_enabled=false`; no X-Accel location in vhost)
Sitemap static serving: **ACTIVE** (nginx try_files + shared files). Index file stale (May 13).
Debug mode: OFF
Debugbar production: NOT PRESENT (packages-dev=0, no debugbar routes)
PRODUCTION_APP_PATH=/opt/www/virtual/SkinbaseNova
→ symlink → /opt/www/virtual/SkinbaseNova.releases/current
→ /opt/www/virtual/SkinbaseNova.releases/releases/20260801-172803-5af95f65-dirty
Local commit: f52879edbb19ecc5d3e807bd4872016f3071dad7 (develop)
Production commit: 5af95f65 (release name 5af95f65-dirty, 2026-08-01)
Same revision: no
Production git: not a git checkout (no .git in release)
Active P0 findings:
QUEUE-001 (90s vs 900s): MITIGATED — Horizon default workers timeout=960
QUEUE rec-job failures: ACTIVE P0 — RecComputeSimilar* fail every night MaxAttemptsExceeded
SITEMAP-001 live-build: NOT the current incident — static files exist
SITEMAP index stale: ACTIVE P1/P0-SEO — sitemap.xml mtime 2026-05-13 while shards refresh daily
STAT-001: NOT currently P0 at 2666 views/24h; design still sync writes (P1)
DB-001: indexes PRESENT; residual UNKNOWN without slow.log. Snapshot table 1.85 GB is the dominant store.
Active P1 findings: swap exhaustion; mail queue unconsumed (LLEN=3); upload 20M vs app 50M; Sentry traces 100%; FPM 10s slowlog volume; CF HTML DYNAMIC; asset TTL 30d not 1y immutable
Major current bottleneck: MySQL row-read volume + 1.85 GB hourly snapshots; nightly rec jobs failing; memory pressure (swap full)
Largest remaining unknown: contents of /var/log/mysql/slow.log (permission)
M1 readiness: YES for planning. Do not implement in M0.
1. Repository State (LOCAL)
branch: develop
HEAD: f52879edbb19ecc5d3e807bd4872016f3071dad7
commit: f52879ed Current state with latest updates
tree: clean at start of this SSH pass (ahead of origin/develop by 1)
No reset/stash/clean on local or production.
2. Production Host (PRODUCTION SERVER)
hostname: server3
FQDN: server3.klevze.si
OS: Debian GNU/Linux 13 (trixie) 13.6
kernel: 6.12.96+deb13-cloud-amd64
arch: x86_64
uptime: 24 days, 16:58 (sampled 2026-08-23 13:43 CEST)
3. Hardware Baseline (PRODUCTION SERVER)
CPU model: AMD EPYC Processor (with IBPB)
sockets: 1
cores: 8
threads: 8 (1 thread/core)
RAM: 23 GiB
swap: 2.0 GiB, ~2.0 GiB used, 2.7 MiB free
disk: /dev/sda 400G → sda1 399.9G on /
filesystem: ~394G, 185G used, 49%
rotational: 0 → SSD/cloud SSD (not HDD)
dmidecode not used (would need sudo).
4. Current System Load (PRODUCTION SERVER)
Short snapshot only (limitation: not peak-hour).
load average: 0.85, 2.47, 2.88
vmstat 5s: idle ~76–83%, wa ~0–1%, r=3–6
CPU notables: mysqld ~20%, containerd/dockerd ~45% each (other workloads on same host), redis ~5%
RAM: 14 GiB used + 7 GiB cache; 8.6 GiB available
swap: fully used
Largest RSS:
| Process | RSS | Note |
|---|---|---|
| mysqld | ~7.7 GiB | 32.7% |
| python uvicorn × several | up to 1.3 GiB | vision/ML stack |
| clamd | ~1.1 GiB | |
| meilisearch | ~397–401 MiB | |
| redis | ~164 MiB | |
| php-fpm skinbase | ~95–100 MiB each | 6 workers sampled |
| node ssr.js | ~122 MiB |
Classification: WATCH / CONCERNING (swap exhausted; shared host with Docker/ML/Qdrant/Gitea/Crowdsec). Not a Skinbase-only box.
5. Process Inventory (PRODUCTION SERVER)
| Component | Reality |
|---|---|
| Horizon | YES — Supervisor skinbase-horizon, pid running ~1d |
standalone queue:work |
NO |
| Supervisor | YES — horizon + ssr |
| systemd queue workers | NO |
| Inertia SSR | YES — node bootstrap/ssr/ssr.js as www-data |
| Reverb | YES — systemd skinbase-reverb.service, 127.0.0.1:8080 |
| Meilisearch | YES — local systemd |
| Redis | local 127.0.0.1:6379 |
| MySQL | local /usr/sbin/mysqld |
| PHP-FPM | php8.4, pool skinbase + unused pool www |
| nginx | master + 8 workers, www-data |
| Docker/ML | uvicorn:8000, qdrant, gitea — co-tenant |
6. nginx (PRODUCTION SERVER)
version: nginx/1.26.3
vhost: /etc/nginx/sites-enabled/skinbase.org.conf
root: /opt/www/virtual/SkinbaseNova/public
listen: 80 → 301 HTTPS; 443 ssl; http2 on
PHP: unix:/run/php/php8.4-fpm-skinbase.sock
client_max_body_size: 24m
gzip: on (nginx.conf)
brotli: on (nginx.conf)
real IP: /etc/nginx/conf.d/00-cloudflare-realip.conf ACTIVE
search: location = /search + conf.d/13-skinbase-search-protection.conf (20r/m, abusive query map)
HTTP/3: not in this vhost (Cloudflare alt-svc h3 is edge-only)
Snippet vs live:
| Snippet | Status |
|---|---|
sitemaps.conf (max-age=21600, try_files) |
ACTIVE (inlined in vhost) |
static-cache.conf 1y immutable /build/assets |
NOT ACTIVE — generic expires 30d for css/js/images |
| download-accel.conf | NOT ACTIVE |
| search-rate-limit.conf (repo) | EQUIVALENT ACTIVE as 13-skinbase-search-protection.conf |
| upstream-error-pages.conf | NOT seen in vhost |
Bug/risk: sitemap locations use try_files $uri @php but no location @php exists in the vhost. Missing files will not fall through to Laravel as the snippet comments claim.
No nginx -T via sudo (password required). Readable vhost was enough.
7–9. HTTP / CDN (PUBLIC CLOUDFLARE/HTTP + PRODUCTION)
Unchanged from edge probes, now correlated with origin:
- Homepage Cache-Control matches Laravel guest headers; CF DYNAMIC
- Build assets CF HIT, origin
expires 30d - CDN webp
max-age=31536000HIT + Polish - Brotli CONFIRMED at edge; origin brotli on
- HTTP/2 CONFIRMED at origin (
http2 on); this Windows curl was HTTP/1.1 to CF - HTTP/3 offered by CF (
alt-svc), not configured on origin vhost
10–12. PHP / OPcache / FPM (PRODUCTION SERVER)
PHP: 8.4.24
FPM master: /etc/php/8.4/fpm/php-fpm.conf
pool: [skinbase] user=skinbase
listen: /run/php/php8.4-fpm-skinbase.sock
pm: dynamic
max_children: 14
start_servers: 4
min/max spare: 2 / 6
max_requests: 500
request_terminate_timeout: 120s
slowlog: 10s → /var/log/php8.4-fpm-skinbase-slow.log (23 MB, 139556 lines)
memory_limit: 256M (pool)
max_execution: 60s (pool)
upload/post: 20M / 24M (pool) ← app allows 50M images / 200M archives
CLI upload_max: 2M (irrelevant to FPM)
extensions: opcache, redis, pcntl, posix, intl, mbstring, curl, gd, pdo_mysql
imagick: NOT in php -m
opcache (FPM php.ini): memory 256MB, interned 16, max_files 20000, validate_timestamps=1, revalidate_freq=2, jit=off
opcache runtime hit rate: UNKNOWN (no status scrape of /fpm-status from remote; localhost-only)
Workers sampled: 6 skinbase processes, avg RSS ~95 MB, total ~570 MB.
Theoretical ceiling: 14 × 256 MB = 3.6 GB if every worker hit memory_limit. Typical: 14 × 95 MB ≈ 1.3 GB.
Available RAM 8.6 GiB but swap already full — do not raise max_children in M0.
13–15. Laravel Production Configuration (PRODUCTION SERVER)
php artisan about + tinker config() (no .env dump):
APP_ENV=production
APP_DEBUG=false
CACHE_STORE=redis
SESSION_DRIVER=redis
QUEUE_CONNECTION=redis
SCOUT_DRIVER=meilisearch
FILESYSTEM_DISK=local
UPLOAD_QUEUE_DERIVATIVES=false
DOWNLOAD_ACCEL_ENABLED=false
DOWNLOAD_ACCEL_PATH=/internal/originals (path set, flag off, nginx location missing)
SITEMAPS_BUILD_ON_REQUEST=true
SITEMAPS_FALLBACK_TO_LIVE_BUILD=true
SITEMAPS_PREGENERATED_ENABLED=true
SITEMAPS_PREFER_PUBLISHED_RELEASE=true
SITEMAPS_STATIC_PUBLISH_ENABLED=true
REDIS_CLIENT=predis
homepage.cache_store=homepage
homepage.guest_payload_ttl_seconds=1800
recommendations.queue=default
vision.queue=default
discovery.queue=default
broadcasting=reverb
horizon.path=horizon
Sentry enabled, sample rate errors 100%, performance 100%
Laravel caches: config, events, routes, views ALL CACHED.
Debugbar/Telescope/Clockwork routes: none. Composer packages-dev=0 → --no-dev install.
16–22. MySQL (PRODUCTION SERVER)
location: local mysqld
version: 8.4.11-11
database: skinbase
size: 2758.2 MB (data 1230.4 + indexes 1527.8)
buffer pool: 10 GB, 10 instances
max_conn: 120 (Max_used 25)
slow_query_log: ON, long_query_time=0.5s, file /var/log/mysql/slow.log (not readable here)
Slow_queries: 11342 in Uptime 97330s (~27h) ≈ 0.12/s
Threads: connected 3, running 2
InnoDB hit: 1 - 103791/18420885623 ≈ 99.999%
Rows read: 9.70e9 in ~27h (high — job/scan load)
tmp tables: 455567 created, 21 on disk
Largest tables (information_schema estimates):
| Table | ~rows | total MB |
|---|---|---|
| artwork_metric_snapshots_hourly | 8.62M | 1855.5 |
| artwork_downloads | 579k | 167.3 |
| forum_bot_logs | 267k | 141.8 |
| artworks | 50k | 94.2 |
| user_activities | 240k | 69.3 |
| artwork_comments | 172k | 62.2 |
| artwork_view_events | 161k | 25.4 |
| rank_artwork_scores | 50k | 22.4 |
| rec_item_pairs | 74k | 16.0 |
| artwork_stats | 49k | 15.2 |
cache/sessions tables nearly empty (drivers are Redis). jobs table empty (Redis queues).
Indexes (batch1) — PRODUCTION
| Index | Status |
|---|---|
| artworks.idx_public_approved_published_id | PRESENT |
| artworks.idx_public_approved_user_id | PRESENT |
| artworks FULLTEXT title+description | PRESENT |
| snapshots.idx_bucket_artwork | PRESENT |
| rank_artwork_scores idx_mv_trending/new_hot/best | PRESENT |
| tags.artworks_count + idx_tags_artworks_count | PRESENT |
DB-001 re-evaluation
Classification: UNKNOWN as “still 78%”, NOT “unindexed public scans”
April indexes are deployed. Dominant storage is hourly snapshots ≈ 50k artworks × 24h × 7d. Heat/ranking jobs that join this table can still dominate Innodb_rows_read even with idx_bucket_artwork. Cannot confirm fingerprints without slow.log / performance_schema (denied).
Do not add more indexes in M0.
23–25. Redis / Cache / Sessions (PRODUCTION SERVER)
redis_version: 8.10.1
used: 877.55 MB (peak 917 MB)
maxmemory: 2.00 G, policy allkeys-lru
evicted_keys: 0
clients: 20
ops/sec: 44 (instant)
keyspace hits/misses: 1.40M / 1.30M (~52% hit)
db0: 1932 keys; db1: 6641; db2: 55176 (all with TTL)
| System | Redis in production? |
|---|---|
| Cache default | YES |
Homepage store homepage |
configured (failover redis→database) |
| Sessions | YES |
| Queues / Horizon | YES |
| Presence | code uses Redis (not separately counted) |
| Stats deltas | flush command scheduled; view path still MySQL sync |
CACHE-001 from the code audit (database cache default) is NOT ACTIVE in production.
26–32. Horizon / Queues / QUEUE-001 (PRODUCTION SERVER)
Horizon running. Supervisors:
| Supervisor | queues | timeout | maxProcesses |
|---|---|---|---|
| supervisor-default | search, default | 960 | 5 |
| supervisor-messaging | broadcasts, notifications | 90 | 3 |
recommendations.queue / vision.queue / discovery.queue = default → consumed by 960s workers.
Standalone supervisor queue:work --timeout=90 from deploy/supervisor/skinbase-queue.conf is NOT the production worker.
Failed jobs: 329 rows (queue:failed listed many; table_rows estimate 159). Dominant:
RecComputeSimilarByBehaviorJob ~02:18 daily MaxAttemptsExceededException
RecComputeSimilarHybridJob ~02:31 daily MaxAttemptsExceededException
Horizon worker tries=1, job $timeout=900, worker timeout 960. Failures are not explained by the old 90s worker. Likely 900s job timeout, 128 MB worker memory, or lock/overlap. Do not retry in M0.
Queue LLEN (Redis): all listed queues 0 except queues:mail = 3. Horizon does not listen to mail. Mail jobs can stall.
QUEUE-001 (90 vs 900): MITIGATED
Nightly rec job failures: ACTIVE P0 (related, different mechanism)
mail queue unconsumed: P1
31–32. Scheduler (PRODUCTION SERVER)
php artisan schedule:list matches routes/console.php (generate sitemaps 10:30/22:30, rec jobs 02:00–02:30, etc.).
How schedule:run is invoked: user crontab empty; sudo crontab not readable; no systemd timer named laravel/schedule. Empirically running (sitemap shards 10:30 today; rec jobs fail 02:18/02:31 daily). Likely root crontab. Do not change.
02:00–05:00 window (code + evidence it fires): rec jobs, ranking/heat, analytics, prune snapshots, sitemap validate — plus nightly rec failures.
33. Meilisearch (PRODUCTION SERVER)
process: /usr/local/bin/meilisearch --config-file-path /etc/meilisearch.toml
health: {"status":"available"}
version/stats: 401 without key (not printed)
RSS: ~401 MB
No reindex.
34–36. Storage / Downloads / Sitemaps
FILESYSTEM_DISK=local; CDN for public derivatives.- Download accel off; originals would stream via PHP if used.
- Sitemaps live under shared
.../shared/public/sitemaps/(20M, 27 xml files). - Shards + academy + users + forum-threads mtime 2026-08-23 10:30–10:31
sitemap.xmlmtime 2026-05-13 21:39, 1906 bytes — generate/publish does not update the index file.- May index omits academy-* families that now exist on disk.
SITEMAP-001 live-build-on-missing: files exist, so crawlers are not building XML in PHP for the index. @php named location missing anyway.
SITEMAP-001 original (PHP live-build stampede): NOT ACTIVE right now
Stale sitemap index vs daily shards: ACTIVE (SEO/ops)
37–38. Views / Downloads
artwork_view_events last 1h: 72
artwork_view_events last 24h: 2666
artwork_downloads last 24h: 4718
view_events table: ~161k rows, 25.4 MB
downloads table: ~579k rows, 167 MB
STAT-001 architecture (sync INSERT+UPDATE, defer=false) is still the code on this release. At ~0.03 views/s it is not the current DB bottleneck. Classify P1 (architecture / growth), not active P0 load.
Downloads were not fetched (would write).
39–47. Latency / Traffic / Resources
Origin access-log QPS not aggregated (would include IPs — skipped). Edge n=1 TTFBs from earlier remain PUBLIC evidence, not origin APM.
FPM slowlog: 139k historical lines, last file mtime Aug 23 05:05, timeout 10s. Scripts are index.php only (no URI dump here).
Resource baseline: see §4. Swap full = CONCERNING.
48. Lighthouse / CWV
Not re-run. Prior lab file is skinbase.top 2026-03-23 — do not use as this baseline.
49. Observability
- Sentry on, 100% error and performance sample (P1 cost).
- Netdata on host.
- Horizon dashboard path
horizon(auth not verified; do not probe). /fpm-statuslocalhost only./stats/basic-auth (not accessed).- MySQL slow.log exists but unreadable to this SSH user.
50. Re-evaluated P0 Findings
| ID | Verdict | Why |
|---|---|---|
| QUEUE-001 90s worker vs 900s job | MITIGATED | Horizon default timeout 960; rec queues aliased to default |
| RecCompute* nightly MaxAttemptsExceeded | ACTIVE P0 | every night 02:18/02:31 on redis@default; 329 failed jobs |
| STAT-001 sync view writes | P1 at current 2666/day | still sync in config/code; not the load leader |
| SITEMAP-001 request-time build | NOT ACTIVE | static files present; @php missing |
| Stale sitemap.xml (May vs Aug shards) | ACTIVE P1 (SEO) | generate writes children, not index |
| DB-001 unindexed aggregates | MITIGATED for missing indexes | batch1 PRESENT. Residual scan cost UNKNOWN; snapshots 1.85 GB |
51. Re-evaluated P1 Findings
- Production Redis cache/session (repo default database is wrong for prod).
- Upload 20M/24M FPM vs app 50M/200M.
UPLOAD_QUEUE_DERIVATIVES=false— sync image work on publish.mailRedis queue not consumed by Horizon.- Sentry performance sampling 100%.
- CF HTML DYNAMIC despite public Cache-Control.
- Asset TTL 30d not 1y immutable.
- Swap 2G full; shared ML/Docker host.
- FPM slowlog 10s historically large.
- Local git newer than production release (
f52879edvs5af95f65). - Session cookies on
/searchand 404 (code-consistent). - Duplicate security headers (app middleware + nginx snippet).
52. Baseline Metric Table
| Metric | Value | Source |
|---|---|---|
| DB size | 2758 MB | information_schema |
| Snapshots table | 1855 MB / ~8.6M rows | information_schema |
| Artworks | ~50k / 94 MB | information_schema |
| Views 24h | 2666 | count on 161k table |
| Downloads 24h | 4718 | count |
| Slow_queries / 27h | 11342 | SHOW STATUS |
| InnoDB pool hit | ~100% | STATUS |
| Redis used | 878 MB / 2 GB | INFO |
| Redis hit ratio | ~52% | INFO |
| FPM workers | 6 / max 14, ~95 MB RSS | ps |
| Horizon | running, timeout 960/90 | ps + artisan |
| Failed jobs | 329 | DB |
| Homepage TTFB (edge) | 146 ms n=1 | PUBLIC |
| Production release | 20260801-172803-5af95f65-dirty | filesystem |
53. Production Risk Matrix
| Risk | Evidence | Impact | Confidence |
|---|---|---|---|
| Nightly rec jobs never succeed | failed_jobs every day | stale similar-art | high |
| Hourly snapshot table 1.85 GB | information_schema | heat/rank I/O | high |
| Swap full + co-tenant ML | free/ps | latency spikes | high |
| Stale sitemap index | mtime May 13 vs shards today | crawl/SEO | high |
| mail queue not consumed | LLEN=3, Horizon queues | delayed mail | high |
| Missing @php location | nginx vhost | missing shard → error not Laravel | high |
| Cannot read slow.log | permissions | DB-001 residual | high |
54. Optimization Targets (do not implement now)
- Rec job failures (timeout/memory/overlap) — after measuring why MaxAttemptsExceeded.
- Snapshot retention / heat query EXPLAIN once slow.log is readable.
- Publish
sitemap.xmlindex with shards. - Consume
mailor stop using that queue name. - Download X-Accel if originals still hit FPM.
- Align upload limits.
- Memory/swap / co-tenancy.
- Sentry sample rates.
- CF HTML cache policy (product decision).
55. M1–M4 Readiness
M1 planning: READY (production drivers known; P0 recast)
M1 implementation: not this milestone
Blockers for DB-001 closeout: mysql slow.log or GRANTs on performance_schema
56. Unknowns
- Root crontab contents (scheduler invocation path)
- MySQL slow.log fingerprints now
- OPcache hit rate (FPM status localhost-only)
- Meilisearch document counts (auth)
- Peak QPS / access-log (IPs omitted on purpose)
- Why rec jobs MaxAttemptsExceeded (timeout vs memory vs exception) — do not dump payloads
- HTTP/3 to origin (not configured)
- iostat (command missing)
57. Repository Changes
Created earlier:
- scripts/collect-production-baseline.sh
Updated:
- docs/skinbase-production-baseline.md
Modified application:
- none
Database changes:
- none
Server configuration changes:
- none
Service restarts:
- none
Cache clears:
- none
Queue operations:
- none (failed jobs listed only, not retried)