Files
SkinbaseNova/docs/optimization-m9-scheduler-overlap.md
T
klevze 8a80aae21e Ship production optimization M1-M12.5A: queues, metrics, HTTP observability, and vector search reliability.
Keep similar-ai from tripping the global circuit on a lone URL 502, clamp Qdrant search to 100, and add Server-Timing plus slow-request logging. Studio shared props, Academy S3 exists caching, heat chunking, and Redis/scheduler hygiene stay in this rollout.
2026-08-25 07:58:47 +02:00

8.3 KiB
Raw Blame History

M9 — Scheduler, Cron Overlap & Background Job Audit

STATUS: COMPLETE (audit + small local schedule hygiene)
PRODUCTION: unchanged (no cron/systemd/Supervisor edits)
WHEN: 2026-08-23 20:23–20:25 CEST

1. Scheduler execution mechanism

Laravel schedule is defined only in routes/console.php (bootstrap/app.php commands:). Kernel::schedule() is empty to avoid double registration.

How schedule:run is invoked: empirically running (live schedule:finish wrapper around leaderboards:refresh; health:tick Redis key present). User crontab empty. /etc/cron.d has no artisan schedule:run. systemd has no schedule timer. Inferred: root crontab (sudo crontab -l not readable by this user). Do not change.

Supervisor runs Horizon + SSR only, not schedule:work.


2. Duplicate-runner check

Check Result
schedule:work processes none
Second schedule:run host single host server3
Duplicate cron.d artisan none
Kernel + console.php double register Kernel schedule empty; schedule:list unique
Stale schedule:work none

NO ISSUE for duplicate runners.


3. Laravel schedule inventory (production schedule:list)

Gates: collections dispatch false, forum AI/sec false, presence prune true, security-report true.

Forum async and collections:dispatch-maintenance absent from list. Presence prune present.

Every-minute: posts/artworks/news/nova-cards publish, health:tick.

Heavy hourly: metrics :02, rankings :07/:37, heat :09/:24/:39/:54, rank-build :15, forum:scan-posts :17, leaderboards :21 (mutex held, >2 min observed), search-reconcile :28, presence prune :33, horizon snapshot :45.

Nightly rec: tags 02:00, behavior 02:15, hybrid 02:30 (queue jobs; M1/M7.1).

Sitemaps: generate 10:30 and 22:30; publish 8 */6; validate 04:45; cleanup 05:00.

Full table: production schedule:list output captured 20:23 CEST (44 events). Local adds overlap names on analytics/uploads; flush minutes 18 instead of 21.

uploads:cleanup previously had no withoutOverlapping (fixed locally). Four analytics:aggregate-* same (fixed locally).

Default overlap expiry: Laravel 1440 minutes except prune-metric 120, publish-scheduled 2, presence prune 25.


4. System cron/timers relevant to Skinbase

Source What
/etc/cron.d/skinbase-public-index-integrity */5 root integrity script
/etc/cron.d/php sessionclean 09,39 (systemd timer wins)
certbot.timer ~22:26
logrotate.timer ~00:02
apt-daily.timer ~02:12
lynis, crowdsec-hub, dpkg-db-backup ~00:00–00:26
ClamAV daemon, not a Skinbase timer

5–6. Timeline / collisions

When Tasks Rank
:21 flush-redis-stats and leaderboards:refresh (refresh held mutex >2 min) MEDIUM — local: flush moved to :18
:33 collections:sync-lifecycle (DB) + presence prune (Redis) LOW — different resources
:15 rank-build-lists + homepage warm LOW
02:00–02:30 rec tags→behavior→hybrid on default queue; apt-daily ~02:12 MEDIUM (queue contention, not schedule mutex). Offsets are guessed, not a completion barrier. Do not redesign in M9.
03:00–04:45 uploads, analytics, reset-windowed, enhance, academy prune, metric prune, sitemap validate LOW–MEDIUM overnight cluster, sequential minutes
03:30 reset-windowed-stats-24h and enhance:cleanup LOW (own mutexes, enhance background)
every minute 5 commands NO ISSUE if each is cheap

7. Measured runtimes

Job Evidence
leaderboards:refresh process etime 2m23s still running; schedule mutex listed
IndexUserJob / MakeSearchable Horizon 16–108 ms
sitemaps generate prior ~65s; public/sitemaps/sitemap.xml mtime 14:51 (release/publish, not 10:30)
presence prune first scheduled run ~81k members (M8 verify); hourlyAt(33) keep
rec nightly historical MaxAttempts 02:18 / 02:31 (M1; mitigated in code, not schedule)
health:tick Redis SETEX; key 1787509506 present

Did not run expensive commands for benchmarks.


8. withoutOverlapping

Missing (fixed locally): uploads:cleanup, four analytics aggregates.

TTL 1440m default: crash suppresses a daily job up to 24h — acceptable for daily. Presence 25m matches catch-up batches. Metric prune 120m. Publish 2m.

leaderboards uses default 1440 while runtime ~2–3 min — OK.


9. runInBackground

Used for publish, rankings, metrics, heat, sitemaps, forum scan-posts, presence prune, leaderboards, etc.

Observed: leaderboards:refresh spawned via schedule:finish and stdout to /dev/null. Failures may not fail schedule:run. Do not remove globally. Mutex still applied (Has Mutex on list).


10. Every-minute

health:tick: one Redis SETEX, TTL 300. NO CHANGE.

Publish-scheduled ×3 + posts: withoutOverlapping(2) / default. Keep.

posts:warm-trending odd minutes: overlap protected.


11. Recommendation

02:00 tags, 02:15 behavior, 02:30 hybrid, everyFourHours pairs at :00. M1/M7.1 job code intact. No completion wait between stages — hybrid may run while tag batches still on default. Historical failures were retry_after (M1), not this offset. No architecture change.


12. Ranking / leaderboards

nova:recalculate-rankings :07/:37 background. rank-build-lists :15 (dispatches unique scopes, 6h lock). leaderboards:refresh :21 DB-heavy, >2 min. Duplicate scope stacking addressed in M7.1. NO further frequency change.


13. Metrics

Snapshot :02 (~50k upserts). Prune 04:25 overlap 120 — keep (outside :02). Studio 30d is HTTP not schedule.


14. Presence

Enabled on production. hourlyAt(33), overlap 25. Redis-heavy; :33 vs lifecycle DB. Keep. Do not accelerate.


15. Forum / collections

forum:scan-posts --limit=250 hourly :17, in-process, overlap, background. Async forum absent. Collections dispatch absent. collections:sync-lifecycle every 10 min (3,13,23,33,43,53) — invite/lifecycle SQL, not the orphan queue. NO CHANGE to frequency without row-count evidence.


16. Sitemap

generate 10:30/22:30 background overlap. publish every 6h at :08. validate 04:45. cleanup 05:00. Two daily generates remain justified (fresh shards). Do not regress M3.


17. Horizon snapshot

hourlyAt(45). Current prefix skinbase_horizon:. Horizon log shows live job timings — metrics path works. NO CHANGE.


18. OS maintenance

apt-daily ~02:12 vs rec 02:00–02:30 (MEDIUM IO). logrotate ~00:02. Do not edit OS timers.


19. Deploy

sync.sh → deploy-production.sh, restarts Horizon/SSR. Background schedule:finish children can outlive symlink (leaderboards pattern). Mutex in cache/Redis (not files found on disk). No deploy script change — no proven double scheduler on release switch.


20. Ranked findings

Sev Finding Action
MEDIUM flush-redis-stats and leaderboards both :21; leaderboards >2 min Local: flush :18
LOW uploads:cleanup / analytics aggregates lacked withoutOverlapping Local: added
MEDIUM rec stages overlap on default queue by clock Document only
LOW :33 lifecycle + presence Different stores; keep
NO CHANGE health:tick, presence hour, metric retention, forum/collections gates, sitemap times, Horizon snapshot, worker counts

Task Current Recommended
flush-redis-stats 1,11,21,31,41,51 1,11,18,31,41,51
uploads:cleanup daily 03:00, no overlap + withoutOverlapping + name
analytics aggregates daily, no overlap + withoutOverlapping + names

Rollback: revert those three edits in routes/console.php.


22–25. Files / tests / deploy

routes/console.php, docs/optimization-m9-scheduler-overlap.md, tests/Feature/Scheduler/ScheduleInventoryTest.php.

Deploy: ship routes/console.php; no cron edit; no Horizon restart required for schedule (next schedule:run loads new events). Rollback: revert file.

Do not increase workers/FPM/Redis/MySQL/swap or re-enable orphan queues.