2e6a981120
release / release (push) Successful in 13s
M4 — health report (the 0.0.4 CHANGELOG entry, folded into this release):
- core/health.py: scan journalctl (Xid/panic/OOM/MCE/AER/thermal), SMART,
NVIDIA driver mismatch, journald persistence, live temps -> findings
- CLI `rigdoctor report` (text/JSON); GUI Health tab; scanner tests
M9 — installer (first cut):
- core/{catalog,sysenv,installer}.py; `rigdoctor install [--check] [-y]`
- GUI Setup tab: detect distro/GPU, show optional components, one-click
install of missing apt packages via pkexec/sudo
M13 — update check (check half):
- core/updates.py; sidebar shows up-to-date / "Update to v…" / unavailable
Plus tests, version bump to 0.0.5, CHANGELOG, and doc status updates.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
3.9 KiB
3.9 KiB
RigDoctor — Roadmap (DRAFT v0.2)
Phased so the seed use case (capturing the RTX 3070 crash / black-screen events) is solved
early, before the broader "tool for all Linux gamers" work. Stack: Python 3 + Qt/PySide6;
Ubuntu + NVIDIA first; .deb distribution (see DECISIONS.md).
Phase 0 — Workspace & spec (done)
- Create repo + docs scaffold
- Settle the foundational decisions D1–D11 (name, language, platform/GPU priority, MVP scope, trigger model, packaging, scope-of-action, GUI/tray)
- Lock the MVP scope (M1 + M3 + M4, NVIDIA-only)
Phase 1 — MVP: capture this crash (Essential bundle, NVIDIA-only, CLI)
- M1 sensor core (NVIDIA via nvidia-smi + hwmon for CPU/RAM/NVMe), stdlib-only
- M3 crash-capture logger (JSONL, fsync per sample, GPU-lost detection, size rotation)
- Manual trigger mode (
rigdoctor record run/start/stop/status);systemd --userservice + other trigger modes in Phase 4 (runis already the service entrypoint) - M4 health report (Xid/panic/OOM/MCE/AER/thermal scan + SMART + driver-mismatch + journald-persistence + live temps, suggested fixes only — D9; GPU-firmware verify deferred)
record reportpost-crash summary (peak temps/power per subsystem, events, last N samples)- Exit criteria: user can run it during gaming and, after a freeze/black-screen, see the last readings + a plausible cause.
Phase 2 — Live monitor (terminal)
- M2 TUI dashboard (current/min/max, grouped, throttle highlighting)
- M8 basic alerting (overheat/throttle/GPU-lost notifications)
Phase 3 — Diagnostics breadth
- M5 system inventory + exportable report
- M6 gaming environment checks (suggest-only)
- SMART integration (smartmontools if present)
Phase 4 — Desktop UI & installer
- M10 desktop GUI (PySide6: dashboard, log browser, report viewer, logger controls)
- M11 tray / menu-bar applet (QSystemTrayIcon: live M1 readouts + Run Diagnostic + supporting actions — D13)
- Guided diagnostic session (pick game → focused M3 capture → M4 scan → findings), shared by tray/GUI/CLI
- Logger trigger modes: always-on + game-launch (D12 — wrapper first:
rigdoctor wrap %command%+ global Steam compat-tool; zero-config watcher (Steam RunningAppID + /proc) and GameMode hook follow) - [~] M9 interactive installer — done: distro/GPU detection + optional-dependency install
(
rigdoctor install, GUI Setup tab). Pending: module-selection config +systemd --userservice enable + trigger-mode pick. .debpackaging (D8) declaring per-bundle deps incl. python3-pyside6 for Desktop UI
Phase 5 — Breadth (later)
- AMD GPU support in M1 (Steam Deck / Radeon)
- Intel GPU best-effort
- [~] M13 auto-update (D18) — done: launch-time version check shown in the GUI sidebar
(up-to-date / "Update to v…" / unavailable). Pending: no-root self-update of the
user-local install from the public Gitea releases;
rigdoctor update. - (Later, separate milestone) Optional auto-apply of suggested fixes behind explicit consent — currently out of scope (D9)
Phase 6 — Session sharing / remote assist (M12, D16)
Escalating ladder, built in order:
- Tier 1:
share export— diagnostic bundle (inventory + recent log + report); B opens it in RigDoctor. One-way, safest. - Tier 2: live read-only view (local server + user-chosen tunnel: Tailscale/cloudflared/ SSH; no hosted relay), token-gated, A approves, revocable.
- Tier 3: gated interactive terminal (wrap tmate/sshx; read-only default, read-write on explicit consent), with session audit log.
Out of scope: stress/repro module (D7); multi-distro support and packaging beyond Ubuntu/apt +
.deb(D15) — a thin seam is kept but not built out.
Dropped: stress / repro module (D7) — not on the roadmap.