Files
SkinbaseNova/docs/skinbase-production-baseline.md
klevze 8a80aae21e Ship production optimization M1-M12.5A: queues, metrics, HTTP observability, and vector search reliability.
Keep similar-ai from tripping the global circuit on a lone URL 502, clamp Qdrant search to 100, and add Server-Timing plus slow-request logging. Studio shared props, Academy S3 exists caching, heat chunking, and Redis/scheduler hygiene stay in this rollout.
2026-08-25 07:58:47 +02:00

21 KiB
Raw Permalink Blame History

Skinbase.org — M0 Production Baseline & Runtime Audit

STATUS: COMPLETE

Primary origin runtime unknowns were resolved via read-only ssh server3. Some items remain UNKNOWN (MySQL slow.log not readable by this user; root crontab not readable without sudo; performance_schema denied to skinbase@localhost; iostat not installed).

Nothing was modified on server3. No restarts, no cache clears, no SQL writes, no deploys.

Evidence classes used below:

LOCAL CODE EVIDENCE
PRODUCTION SERVER EVIDENCE
PUBLIC CLOUDFLARE/HTTP EVIDENCE

Executive Summary

Production OS:                 Debian 13 (trixie), kernel 6.12.96+deb13-cloud-amd64, x86_64
CPU:                           AMD EPYC, 8 cores / 8 threads (1 socket)
RAM:                           23 GiB; ~8.6 GiB available; SWAP 2.0 GiB **almost fully used**
Storage:                       400G virtio disk `/dev/sda1` ext4 ~185G/394G (49%); rotational=0 (SSD/NVMe-like)
nginx:                         1.26.3; HTTP/2 on; gzip+brotli in nginx.conf; vhost skinbase.org.conf
PHP:                           8.4.24 (FPM + CLI)
PHP-FPM:                       pool `skinbase`, dynamic, max_children=14, memory_limit=256M, terminate=120s
MySQL:                         8.4.11-11 local; DB `skinbase` 2758 MB; buffer pool 10 GB
Redis:                         8.10.1 local 127.0.0.1:6379; 878 MB used / 2 GB max; allkeys-lru
Meilisearch:                   local systemd; health available; RSS ~401 MB
Queue manager:                 Supervisor → `php artisan horizon` (NOT standalone queue:work)
Cache store:                   **redis** (PRODUCTION; not repo default database)
Session store:                 **redis**
CDN:                           Cloudflare + cdn.skinbase.org RGW/S3
Download acceleration:         **OFF** (`download_accel_enabled=false`; no X-Accel location in vhost)
Sitemap static serving:        **ACTIVE** (nginx try_files + shared files). Index file stale (May 13).
Debug mode:                    OFF
Debugbar production:           NOT PRESENT (packages-dev=0, no debugbar routes)
PRODUCTION_APP_PATH=/opt/www/virtual/SkinbaseNova
  → symlink → /opt/www/virtual/SkinbaseNova.releases/current
  → /opt/www/virtual/SkinbaseNova.releases/releases/20260801-172803-5af95f65-dirty

Local commit:       f52879edbb19ecc5d3e807bd4872016f3071dad7 (develop)
Production commit:  5af95f65 (release name 5af95f65-dirty, 2026-08-01)
Same revision:      no
Production git:     not a git checkout (no .git in release)
Active P0 findings:
  QUEUE-001 (90s vs 900s): MITIGATED — Horizon default workers timeout=960
  QUEUE rec-job failures:  ACTIVE P0 — RecComputeSimilar* fail every night MaxAttemptsExceeded
  SITEMAP-001 live-build:  NOT the current incident — static files exist
  SITEMAP index stale:     ACTIVE P1/P0-SEO — sitemap.xml mtime 2026-05-13 while shards refresh daily
  STAT-001:                NOT currently P0 at 2666 views/24h; design still sync writes (P1)
  DB-001:                  indexes PRESENT; residual UNKNOWN without slow.log. Snapshot table 1.85 GB is the dominant store.

Active P1 findings:        swap exhaustion; mail queue unconsumed (LLEN=3); upload 20M vs app 50M; Sentry traces 100%; FPM 10s slowlog volume; CF HTML DYNAMIC; asset TTL 30d not 1y immutable

Major current bottleneck:  MySQL row-read volume + 1.85 GB hourly snapshots; nightly rec jobs failing; memory pressure (swap full)

Largest remaining unknown: contents of /var/log/mysql/slow.log (permission)

M1 readiness:              YES for planning. Do not implement in M0.

1. Repository State (LOCAL)

branch: develop
HEAD:   f52879edbb19ecc5d3e807bd4872016f3071dad7
commit: f52879ed Current state with latest updates
tree:   clean at start of this SSH pass (ahead of origin/develop by 1)

No reset/stash/clean on local or production.


2. Production Host (PRODUCTION SERVER)

hostname:     server3
FQDN:         server3.klevze.si
OS:           Debian GNU/Linux 13 (trixie) 13.6
kernel:       6.12.96+deb13-cloud-amd64
arch:         x86_64
uptime:       24 days, 16:58 (sampled 2026-08-23 13:43 CEST)

3. Hardware Baseline (PRODUCTION SERVER)

CPU model:     AMD EPYC Processor (with IBPB)
sockets:       1
cores:         8
threads:       8 (1 thread/core)
RAM:           23 GiB
swap:          2.0 GiB, ~2.0 GiB used, 2.7 MiB free
disk:          /dev/sda 400G → sda1 399.9G on /
filesystem:    ~394G, 185G used, 49%
rotational:    0  → SSD/cloud SSD (not HDD)

dmidecode not used (would need sudo).


4. Current System Load (PRODUCTION SERVER)

Short snapshot only (limitation: not peak-hour).

load average:  0.85, 2.47, 2.88
vmstat 5s:     idle ~76–83%, wa ~0–1%, r=3–6
CPU notables:  mysqld ~20%, containerd/dockerd ~45% each (other workloads on same host), redis ~5%
RAM:           14 GiB used + 7 GiB cache; 8.6 GiB available
swap:          fully used

Largest RSS:

Process RSS Note
mysqld ~7.7 GiB 32.7%
python uvicorn × several up to 1.3 GiB vision/ML stack
clamd ~1.1 GiB
meilisearch ~397–401 MiB
redis ~164 MiB
php-fpm skinbase ~95–100 MiB each 6 workers sampled
node ssr.js ~122 MiB

Classification: WATCH / CONCERNING (swap exhausted; shared host with Docker/ML/Qdrant/Gitea/Crowdsec). Not a Skinbase-only box.


5. Process Inventory (PRODUCTION SERVER)

Component Reality
Horizon YES — Supervisor skinbase-horizon, pid running ~1d
standalone queue:work NO
Supervisor YES — horizon + ssr
systemd queue workers NO
Inertia SSR YES — node bootstrap/ssr/ssr.js as www-data
Reverb YES — systemd skinbase-reverb.service, 127.0.0.1:8080
Meilisearch YES — local systemd
Redis local 127.0.0.1:6379
MySQL local /usr/sbin/mysqld
PHP-FPM php8.4, pool skinbase + unused pool www
nginx master + 8 workers, www-data
Docker/ML uvicorn:8000, qdrant, gitea — co-tenant

6. nginx (PRODUCTION SERVER)

version: nginx/1.26.3
vhost:   /etc/nginx/sites-enabled/skinbase.org.conf
root:    /opt/www/virtual/SkinbaseNova/public
listen:  80 → 301 HTTPS; 443 ssl; http2 on
PHP:     unix:/run/php/php8.4-fpm-skinbase.sock
client_max_body_size: 24m
gzip:    on (nginx.conf)
brotli:  on (nginx.conf)
real IP: /etc/nginx/conf.d/00-cloudflare-realip.conf ACTIVE
search:  location = /search + conf.d/13-skinbase-search-protection.conf (20r/m, abusive query map)
HTTP/3:  not in this vhost (Cloudflare alt-svc h3 is edge-only)

Snippet vs live:

Snippet Status
sitemaps.conf (max-age=21600, try_files) ACTIVE (inlined in vhost)
static-cache.conf 1y immutable /build/assets NOT ACTIVE — generic expires 30d for css/js/images
download-accel.conf NOT ACTIVE
search-rate-limit.conf (repo) EQUIVALENT ACTIVE as 13-skinbase-search-protection.conf
upstream-error-pages.conf NOT seen in vhost

Bug/risk: sitemap locations use try_files $uri @php but no location @php exists in the vhost. Missing files will not fall through to Laravel as the snippet comments claim.

No nginx -T via sudo (password required). Readable vhost was enough.


7–9. HTTP / CDN (PUBLIC CLOUDFLARE/HTTP + PRODUCTION)

Unchanged from edge probes, now correlated with origin:

  • Homepage Cache-Control matches Laravel guest headers; CF DYNAMIC
  • Build assets CF HIT, origin expires 30d
  • CDN webp max-age=31536000 HIT + Polish
  • Brotli CONFIRMED at edge; origin brotli on
  • HTTP/2 CONFIRMED at origin (http2 on); this Windows curl was HTTP/1.1 to CF
  • HTTP/3 offered by CF (alt-svc), not configured on origin vhost

10–12. PHP / OPcache / FPM (PRODUCTION SERVER)

PHP:              8.4.24
FPM master:       /etc/php/8.4/fpm/php-fpm.conf
pool:             [skinbase] user=skinbase
listen:           /run/php/php8.4-fpm-skinbase.sock
pm:               dynamic
max_children:     14
start_servers:    4
min/max spare:    2 / 6
max_requests:     500
request_terminate_timeout: 120s
slowlog:          10s → /var/log/php8.4-fpm-skinbase-slow.log (23 MB, 139556 lines)
memory_limit:     256M (pool)
max_execution:    60s (pool)
upload/post:      20M / 24M (pool)  ← app allows 50M images / 200M archives
CLI upload_max:   2M (irrelevant to FPM)
extensions:       opcache, redis, pcntl, posix, intl, mbstring, curl, gd, pdo_mysql
imagick:          NOT in php -m
opcache (FPM php.ini): memory 256MB, interned 16, max_files 20000, validate_timestamps=1, revalidate_freq=2, jit=off
opcache runtime hit rate: UNKNOWN (no status scrape of /fpm-status from remote; localhost-only)

Workers sampled: 6 skinbase processes, avg RSS ~95 MB, total ~570 MB.

Theoretical ceiling: 14 × 256 MB = 3.6 GB if every worker hit memory_limit. Typical: 14 × 95 MB ≈ 1.3 GB.

Available RAM 8.6 GiB but swap already full — do not raise max_children in M0.


13–15. Laravel Production Configuration (PRODUCTION SERVER)

php artisan about + tinker config() (no .env dump):

APP_ENV=production
APP_DEBUG=false
CACHE_STORE=redis
SESSION_DRIVER=redis
QUEUE_CONNECTION=redis
SCOUT_DRIVER=meilisearch
FILESYSTEM_DISK=local
UPLOAD_QUEUE_DERIVATIVES=false
DOWNLOAD_ACCEL_ENABLED=false
DOWNLOAD_ACCEL_PATH=/internal/originals   (path set, flag off, nginx location missing)
SITEMAPS_BUILD_ON_REQUEST=true
SITEMAPS_FALLBACK_TO_LIVE_BUILD=true
SITEMAPS_PREGENERATED_ENABLED=true
SITEMAPS_PREFER_PUBLISHED_RELEASE=true
SITEMAPS_STATIC_PUBLISH_ENABLED=true
REDIS_CLIENT=predis
homepage.cache_store=homepage
homepage.guest_payload_ttl_seconds=1800
recommendations.queue=default
vision.queue=default
discovery.queue=default
broadcasting=reverb
horizon.path=horizon
Sentry enabled, sample rate errors 100%, performance 100%

Laravel caches: config, events, routes, views ALL CACHED.

Debugbar/Telescope/Clockwork routes: none. Composer packages-dev=0 → --no-dev install.


16–22. MySQL (PRODUCTION SERVER)

location:     local mysqld
version:      8.4.11-11
database:     skinbase
size:         2758.2 MB (data 1230.4 + indexes 1527.8)
buffer pool:  10 GB, 10 instances
max_conn:     120 (Max_used 25)
slow_query_log: ON, long_query_time=0.5s, file /var/log/mysql/slow.log (not readable here)
Slow_queries: 11342 in Uptime 97330s (~27h) ≈ 0.12/s
Threads:      connected 3, running 2
InnoDB hit:   1 - 103791/18420885623 ≈ 99.999%
Rows read:    9.70e9 in ~27h  (high — job/scan load)
tmp tables:   455567 created, 21 on disk

Largest tables (information_schema estimates):

Table ~rows total MB
artwork_metric_snapshots_hourly 8.62M 1855.5
artwork_downloads 579k 167.3
forum_bot_logs 267k 141.8
artworks 50k 94.2
user_activities 240k 69.3
artwork_comments 172k 62.2
artwork_view_events 161k 25.4
rank_artwork_scores 50k 22.4
rec_item_pairs 74k 16.0
artwork_stats 49k 15.2

cache/sessions tables nearly empty (drivers are Redis). jobs table empty (Redis queues).

Indexes (batch1) — PRODUCTION

Index Status
artworks.idx_public_approved_published_id PRESENT
artworks.idx_public_approved_user_id PRESENT
artworks FULLTEXT title+description PRESENT
snapshots.idx_bucket_artwork PRESENT
rank_artwork_scores idx_mv_trending/new_hot/best PRESENT
tags.artworks_count + idx_tags_artworks_count PRESENT

DB-001 re-evaluation

Classification: UNKNOWN as “still 78%”, NOT “unindexed public scans”

April indexes are deployed. Dominant storage is hourly snapshots ≈ 50k artworks × 24h × 7d. Heat/ranking jobs that join this table can still dominate Innodb_rows_read even with idx_bucket_artwork. Cannot confirm fingerprints without slow.log / performance_schema (denied).

Do not add more indexes in M0.


23–25. Redis / Cache / Sessions (PRODUCTION SERVER)

redis_version: 8.10.1
used:          877.55 MB (peak 917 MB)
maxmemory:     2.00 G, policy allkeys-lru
evicted_keys:  0
clients:       20
ops/sec:       44 (instant)
keyspace hits/misses: 1.40M / 1.30M  (~52% hit)
db0: 1932 keys; db1: 6641; db2: 55176 (all with TTL)
System Redis in production?
Cache default YES
Homepage store homepage configured (failover redis→database)
Sessions YES
Queues / Horizon YES
Presence code uses Redis (not separately counted)
Stats deltas flush command scheduled; view path still MySQL sync

CACHE-001 from the code audit (database cache default) is NOT ACTIVE in production.


26–32. Horizon / Queues / QUEUE-001 (PRODUCTION SERVER)

Horizon running. Supervisors:

Supervisor queues timeout maxProcesses
supervisor-default search, default 960 5
supervisor-messaging broadcasts, notifications 90 3

recommendations.queue / vision.queue / discovery.queue = default → consumed by 960s workers.

Standalone supervisor queue:work --timeout=90 from deploy/supervisor/skinbase-queue.conf is NOT the production worker.

Failed jobs: 329 rows (queue:failed listed many; table_rows estimate 159). Dominant:

RecComputeSimilarByBehaviorJob  ~02:18 daily  MaxAttemptsExceededException
RecComputeSimilarHybridJob      ~02:31 daily  MaxAttemptsExceededException

Horizon worker tries=1, job $timeout=900, worker timeout 960. Failures are not explained by the old 90s worker. Likely 900s job timeout, 128 MB worker memory, or lock/overlap. Do not retry in M0.

Queue LLEN (Redis): all listed queues 0 except queues:mail = 3. Horizon does not listen to mail. Mail jobs can stall.

QUEUE-001 (90 vs 900): MITIGATED
Nightly rec job failures: ACTIVE P0 (related, different mechanism)
mail queue unconsumed: P1

31–32. Scheduler (PRODUCTION SERVER)

php artisan schedule:list matches routes/console.php (generate sitemaps 10:30/22:30, rec jobs 02:00–02:30, etc.).

How schedule:run is invoked: user crontab empty; sudo crontab not readable; no systemd timer named laravel/schedule. Empirically running (sitemap shards 10:30 today; rec jobs fail 02:18/02:31 daily). Likely root crontab. Do not change.

02:00–05:00 window (code + evidence it fires): rec jobs, ranking/heat, analytics, prune snapshots, sitemap validate — plus nightly rec failures.


33. Meilisearch (PRODUCTION SERVER)

process: /usr/local/bin/meilisearch --config-file-path /etc/meilisearch.toml
health:  {"status":"available"}
version/stats: 401 without key (not printed)
RSS:     ~401 MB

No reindex.


34–36. Storage / Downloads / Sitemaps

  • FILESYSTEM_DISK=local; CDN for public derivatives.
  • Download accel off; originals would stream via PHP if used.
  • Sitemaps live under shared .../shared/public/sitemaps/ (20M, 27 xml files).
  • Shards + academy + users + forum-threads mtime 2026-08-23 10:30–10:31
  • sitemap.xml mtime 2026-05-13 21:39, 1906 bytes — generate/publish does not update the index file.
  • May index omits academy-* families that now exist on disk.

SITEMAP-001 live-build-on-missing: files exist, so crawlers are not building XML in PHP for the index. @php named location missing anyway.

SITEMAP-001 original (PHP live-build stampede): NOT ACTIVE right now
Stale sitemap index vs daily shards: ACTIVE (SEO/ops)

37–38. Views / Downloads

artwork_view_events last 1h:   72
artwork_view_events last 24h:  2666
artwork_downloads last 24h:    4718
view_events table:             ~161k rows, 25.4 MB
downloads table:               ~579k rows, 167 MB

STAT-001 architecture (sync INSERT+UPDATE, defer=false) is still the code on this release. At ~0.03 views/s it is not the current DB bottleneck. Classify P1 (architecture / growth), not active P0 load.

Downloads were not fetched (would write).


39–47. Latency / Traffic / Resources

Origin access-log QPS not aggregated (would include IPs — skipped). Edge n=1 TTFBs from earlier remain PUBLIC evidence, not origin APM.

FPM slowlog: 139k historical lines, last file mtime Aug 23 05:05, timeout 10s. Scripts are index.php only (no URI dump here).

Resource baseline: see §4. Swap full = CONCERNING.


48. Lighthouse / CWV

Not re-run. Prior lab file is skinbase.top 2026-03-23 — do not use as this baseline.


49. Observability

  • Sentry on, 100% error and performance sample (P1 cost).
  • Netdata on host.
  • Horizon dashboard path horizon (auth not verified; do not probe).
  • /fpm-status localhost only.
  • /stats/ basic-auth (not accessed).
  • MySQL slow.log exists but unreadable to this SSH user.

50. Re-evaluated P0 Findings

ID Verdict Why
QUEUE-001 90s worker vs 900s job MITIGATED Horizon default timeout 960; rec queues aliased to default
RecCompute* nightly MaxAttemptsExceeded ACTIVE P0 every night 02:18/02:31 on redis@default; 329 failed jobs
STAT-001 sync view writes P1 at current 2666/day still sync in config/code; not the load leader
SITEMAP-001 request-time build NOT ACTIVE static files present; @php missing
Stale sitemap.xml (May vs Aug shards) ACTIVE P1 (SEO) generate writes children, not index
DB-001 unindexed aggregates MITIGATED for missing indexes batch1 PRESENT. Residual scan cost UNKNOWN; snapshots 1.85 GB

51. Re-evaluated P1 Findings

  • Production Redis cache/session (repo default database is wrong for prod).
  • Upload 20M/24M FPM vs app 50M/200M.
  • UPLOAD_QUEUE_DERIVATIVES=false — sync image work on publish.
  • mail Redis queue not consumed by Horizon.
  • Sentry performance sampling 100%.
  • CF HTML DYNAMIC despite public Cache-Control.
  • Asset TTL 30d not 1y immutable.
  • Swap 2G full; shared ML/Docker host.
  • FPM slowlog 10s historically large.
  • Local git newer than production release (f52879ed vs 5af95f65).
  • Session cookies on /search and 404 (code-consistent).
  • Duplicate security headers (app middleware + nginx snippet).

52. Baseline Metric Table

Metric Value Source
DB size 2758 MB information_schema
Snapshots table 1855 MB / ~8.6M rows information_schema
Artworks ~50k / 94 MB information_schema
Views 24h 2666 count on 161k table
Downloads 24h 4718 count
Slow_queries / 27h 11342 SHOW STATUS
InnoDB pool hit ~100% STATUS
Redis used 878 MB / 2 GB INFO
Redis hit ratio ~52% INFO
FPM workers 6 / max 14, ~95 MB RSS ps
Horizon running, timeout 960/90 ps + artisan
Failed jobs 329 DB
Homepage TTFB (edge) 146 ms n=1 PUBLIC
Production release 20260801-172803-5af95f65-dirty filesystem

53. Production Risk Matrix

Risk Evidence Impact Confidence
Nightly rec jobs never succeed failed_jobs every day stale similar-art high
Hourly snapshot table 1.85 GB information_schema heat/rank I/O high
Swap full + co-tenant ML free/ps latency spikes high
Stale sitemap index mtime May 13 vs shards today crawl/SEO high
mail queue not consumed LLEN=3, Horizon queues delayed mail high
Missing @php location nginx vhost missing shard → error not Laravel high
Cannot read slow.log permissions DB-001 residual high

54. Optimization Targets (do not implement now)

  1. Rec job failures (timeout/memory/overlap) — after measuring why MaxAttemptsExceeded.
  2. Snapshot retention / heat query EXPLAIN once slow.log is readable.
  3. Publish sitemap.xml index with shards.
  4. Consume mail or stop using that queue name.
  5. Download X-Accel if originals still hit FPM.
  6. Align upload limits.
  7. Memory/swap / co-tenancy.
  8. Sentry sample rates.
  9. CF HTML cache policy (product decision).

55. M1–M4 Readiness

M1 planning: READY (production drivers known; P0 recast)
M1 implementation: not this milestone
Blockers for DB-001 closeout: mysql slow.log or GRANTs on performance_schema

56. Unknowns

  • Root crontab contents (scheduler invocation path)
  • MySQL slow.log fingerprints now
  • OPcache hit rate (FPM status localhost-only)
  • Meilisearch document counts (auth)
  • Peak QPS / access-log (IPs omitted on purpose)
  • Why rec jobs MaxAttemptsExceeded (timeout vs memory vs exception) — do not dump payloads
  • HTTP/3 to origin (not configured)
  • iostat (command missing)

57. Repository Changes

Created earlier:
- scripts/collect-production-baseline.sh

Updated:
- docs/skinbase-production-baseline.md

Modified application:
- none

Database changes:
- none

Server configuration changes:
- none

Service restarts:
- none

Cache clears:
- none

Queue operations:
- none (failed jobs listed only, not retried)