Hybrid AI Development Model — Claude Review

Plan review only. Nothing was created, installed, cleaned, deployed or configured. The only file written is this review, in the session scratchpad. Read-only checks were run on the VPS to test the model against what already exists.

What I observed on the VPS (read-only)

Observation Why it matters
The Lucky.ai dev repo /home/claude/dev/luckyai is a git repo (68 commits, branch ci/case-intelligence) with no remote configured The premise "private Git repo is the source of truth" is not true yet. The VPS disk is currently the only copy
Prod /home/claude/luckyai is not a git repo and is owned by claude (755, plus env files with .pre-* backups beside them) The agent user can edit live production. This is exactly the failure the plan says to avoid
/home/claude/case, luckyai-data, cdesktop-runtime are not git repos Unversioned data and code exist next to versioned code
No sudo, no gh, no git credential helper, one SSH private key present in ~/.ssh (I did not check where it points) The agent cannot deploy as root. A remote git host is not set up yet
Lucky.ai dev has README.md, docs/, scripts/ but no AGENTS.md or CLAUDE.md. cdesktop-custom has both files, side by side Agent rules are missing in one project and duplicated in another
Existing scripts: luckyai-dev.sh, backup-now.sh, restore-vps.sh, verify-backup.sh, ci-*.sh. No check/test/build/deploy names, no deploy script A good base to standardise from, but no contract yet
~/.claude/settings.json holds model, status line, plugins. Skills live in ~/.claude/skills and ~/.agent/skills. Claude memory holds user preferences Some working knowledge lives outside Git, in Claude-specific places
Host is 2 GB RAM, 1 vCPU, about 2.0 GB free disk (90% used). Each Claude process is 300–400 MB RSS Budgets must be host-level, not only per project
Lucky.ai already runs Playwright on the VPS (one browser cache is 300–650 MB) The plan's "Playwright → WSL" rule conflicts with current practice
Earlier today /tmp/claude-1000 held about 1.5 GB of agent scratch. 705 MB of it was two Python venvs Agent leftovers, not projects, were the biggest disk cost I found

A. Verdict

PASS WITH CHANGES

The core idea is sound: one Git remote, a light VPS, a heavy WSL, a native Windows, and the agent treated as replaceable. Four things need fixing before it can be called a standard:

  1. Git has no remote yet.
  2. The agent can still write production.
  3. The "where does this run" policy is advisory text, not enforced.
  4. Too much project knowledge still lives in Claude-specific places.

B. Key strengths


C. Problems / risks

High

  1. No Git remote, so no source of truth. The VPS is meant to be "replaceable", yet today it holds the only copy of the dev repo. Fix: create a private remote first, and push from both VPS and WSL. Hosted private repo is the cheapest choice. A self-hosted Gitea would cost RAM, and a bare repo on WSL is unreachable while WSL sleeps.
  2. The agent can write production. Prod is owned by the same user the agent runs as, and there is no sudo. The plan says "AI agents should not casually edit production files", but nothing enforces it. Fix: agent runs as a low-privilege user; prod is owned by root (or a deploy user); the agent can only drop an artifact into a write-only incoming/ directory. Deploy is triggered by a human at first, and later by a systemd path unit that runs a root-owned deploy script. That needs no sudo.
  3. The VPS/WSL/Windows policy is documentation, not enforcement. "The AI agent should not need to guess" is only true if the scripts refuse to run in the wrong place. Fix: each machine declares its role (DEV_ENV=vps|wsl|windows, one line in a profile file), and scripts/ share a small guard that refuses heavy work on the VPS unless explicitly overridden. The guard lives in the repo, so it is agent-independent.
  4. The WSL → VPS handoff is undefined. Step 7 ("problems are reported back") and the artifact flow have no channel. WSL sleeps and sits behind NAT, so it must always initiate. See D for a minimal design.

Medium

  1. pnpm command names assume Node projects. Lucky.ai is Python (FastAPI) plus React. Use scripts/*.sh as the contract; they can call pnpm, pytest or anything else.
  2. Reproducibility across machines is unspecified. "Works on WSL, fails on VPS" is a common trap. Pin Node and Python versions (.nvmrc, .python-version), commit lockfiles, install with --frozen-lockfile.
  3. PROJECT_PROFILE.yaml duplicates AGENTS.md until something reads it. Two files that say the same thing drift.
  4. Deploys and database migrations. A rollback that swaps a symlink does not undo a forward migration. Deploy must take a backup first (Lucky.ai already has backup-now.sh and restore-vps.sh), and migrations must be backward-compatible or gated.
  5. Per-project disk budgets miss the real consumers: package caches, Python venvs (Lucky.ai has three), Playwright browsers, agent scratch, session transcripts (about 229 MB under ~/.claude), backups. See G.
  6. Hidden Claude couplings remain (see H).

Low

  1. Line endings: Windows checkouts can turn shell scripts into CRLF. Add .gitattributes with * text=auto eol=lf.
  2. Android builds do not strictly need Windows. Gradle can run in WSL with the command-line SDK. Windows is needed for the emulator and GUI tools. Keep Windows as an opt-in role for native projects only.
  3. dev as a permanent branch adds merge work for a one-person flow. main plus short-lived branches is enough.

D. Recommended architecture (corrected model)

                private Git remote  (canonical: nobody's disk is)
                 ▲          ▲          ▲
        push/pull│          │          │push/pull
                 │          │          │
   VPS  ─────────┘          │          └───────── WSL
   agent (low-priv user)    │                     heavy build, full tests,
   edit, check, light test  │                     Codex, browser tests
   light build              │                     ── verify report ──► VPS incoming/
        │                   │                     ── artifact ───────► VPS incoming/
        ▼                   │
   incoming/ (write-only)   │            Windows: only when a project needs
        │                                native tools (Android emulator, GUI).
        ▼ (human first; systemd path unit later)  Same Git remote, own checkout.
   deploy.sh (root-owned)
   verify checksum → backup → unpack → migrate → health check → swap → verify → rollback on failure
        ▼
   /opt/<project>/releases/<id> + current   (root-owned; agent cannot write)

Handoff design (minimum, no automation platform)

Which checkout is canonical

The Git remote is canonical. Neither checkout is. The VPS checkout is the agent's working copy, and WSL is a verifier that does not commit to shared branches. WSL-only fixes go on wsl/* branches, as the plan says. This also makes the failure model work: if the VPS agent is unavailable, WSL can keep developing against the same remote.


E. VPS / WSL / Windows responsibility matrix

Task VPS WSL Windows
Agent editing, git commit/push ✅ primary fallback —
check (lint, typecheck) ✅ ✅ —
Light tests (unit, small integration) ✅ ✅ —
Light build ✅ (memory-capped) ✅ —
Full tests, browser tests, stress tests ❌ ✅ —
Full build, large installs ❌ ✅ —
Codex, real local runner tests ❌ ✅ —
Production runtime, deploy ✅ (root-owned) ❌ ❌
Android emulator, Android Studio, GUI tools ❌ ❌ ✅
APK/AAB build ❌ optional (CLI SDK) ✅ default
Secrets per-machine .env per-machine .env per-machine

Exception to record: Lucky.ai already runs headless Playwright on the VPS. Either move it to WSL, or allow "headless shell only, one at a time, memory-capped" as a written exception in that project's AGENTS.md.


F. Recommended standard project contract

Required (every project):

README.md                 what it is, how to run it
AGENTS.md                 the single source of agent rules (see below)
scripts/check.sh          fast, VPS-safe: lint + typecheck (target < 60 s)
scripts/test.sh           takes light|full
scripts/build.sh          takes light|full
.env.example              variable names only
.gitattributes            eol=lf
version pins + lockfile   .nvmrc / .python-version, pnpm-lock / requirements lock

Required only for projects that deploy to the VPS:

scripts/deploy.sh         artifact id in; checksum, backup, unpack, health check, swap, rollback
scripts/rollback.sh
docs/ARCHITECTURE.md      one file is enough

Optional: Makefile as sugar (make is already installed); PROJECT_PROFILE.yaml once a tool consumes it.

Script conventions:

AGENTS.md must contain: purpose; repo and directory boundaries; what runs where (points to the guard, not a second copy of it); test and deploy procedure; forbidden actions (destructive commands, editing prod, committing secrets); operating rules (ask before deleting or publishing; long outputs go to a file); a pointer to docs/DECISIONS.md.
CLAUDE.md, where needed, should be a one-line pointer to AGENTS.md (or import it, as I understand Claude Code allows). Keep one copy. The cdesktop-custom project today keeps two copies that can drift.


G. Resource policy (host-level first)

The binding limit is the host, not any one project. At about 2.0 GB free, a single 700 MB project already uses a third of the headroom.

Host rules

Rule Value
Disk floor keep ≥ 2 GB free and ≤ 85% used. Today's 90% is already over the line
Alert 90% used, or under 1.5 GB free
Disk ledger monthly du per project and per agent area, one line each, in a short file
Agent scratch forbid venvs and node_modules in scratch; agent temp older than 3 days is disposable (this alone reclaimed 705 MB this week)
Package caches trim monthly (~/.npm about 98 MB now)
Browsers one Playwright install per host, never deleted while any session runs
RAM at most 1 interactive agent session + 1 heavy job; builds run with NODE_OPTIONS=--max-old-space-size=768 and nice; consider systemd-run --scope -p MemoryMax=
Swap treat sustained swap use (46% today) as a warning sign, not free memory

Per-project budgets (keep, but they were realistic only for Node projects):

Item Default Add
Source + git < 50 MB
Dev dependencies < 250 MB Python venvs count too (Lucky.ai has three)
Production dependencies < 100 MB keep the last 2 releases only
Build output < 50 MB
Logs < 50 MB journald limit
Browsers not in project budget host-level, one copy
Backups not in project budget host-level, with a retention count

The 300–700 MB per project target is realistic for one or two small projects. It is not a safe multiplier for many projects on this host.


H. Agent-replaceability review

Verdict: the infrastructure can already survive replacing Claude. The knowledge cannot yet.

Infrastructure (git, scripts, deploy, systemd, Caddy) has nothing Claude-specific. What still locks the project to Claude:

Hidden dependency Where it lives Fix
Project decisions and reasoning Claude session transcripts (about 229 MB under ~/.claude) Keep docs/DECISIONS.md: one short entry per non-obvious decision
Working preferences (ask before acting, long output as files) Claude memory Mirror project-relevant ones in AGENTS.md; keep the rest as personal preferences
Skills and plugins ~/.claude/skills, ~/.agent/skills, settings.json Treat as tooling. Anything a project needs must be a script in the repo
Duplicate agent files AGENTS.md and CLAUDE.md both hand-written One real file, one pointer
Model IDs and effort settings env files (.pre-opus55, .pre-sonnet55 copies) and models.conf Keep them in config with a comment, not in habit
Permission behaviour (sessions run with dangerous-mode flags) Claude CLI semantics Enforce with OS user, script guards and root-owned prod, which work for any agent
Lucky.ai's product runtime uses Claude Code runner service This is a product dependency, not a development one. Track separately, behind the adapter that the Remote Dev plan already proposes

Test it, don't assume it: run a cold-start drill. Give a fresh agent (Codex, or Claude with no memory) one small task and only the repository. If it cannot finish using README.md, AGENTS.md and the scripts, the standard is not yet agent-independent. Codex reads AGENTS.md natively, as I understand it. Verify that when you first try it.


I. Remote Dev pilot plan

Use the Remote Dev milestones as the pilot. Each hypothesis has a pass/fail check.

Hypothesis Check When
H1: the light path fits the VPS check under 60 s, test light and build light stay under the 768 MB cap, host free disk does not drop below the floor M1
H2: Git is a sufficient boundary Private remote exists; VPS pushes, WSL pulls the same commit, and a wsl/* fix merges back cleanly M1 (needs the remote first)
H3: WSL catches what VPS cannot Phase 1b smoke test: at least one real difference is found and recorded (paths, sleep, CLI behaviour) Phase 1b
H4: the handoff is one command each side wsl-verify.sh on WSL, then the agent reads report-<commit>.md. Count the manual steps; the goal is 1 human action Phase 1b–2
H5: prod is isolated from the agent Agent user cannot write /opt/remote-dev. Deploy succeeds through incoming/, a forced failure rolls back Phase 5
H6: an agent can be replaced Cold-start drill with a second agent, repo only after M1
H7: resource policy holds Monthly ledger: disk, RAM peaks, scratch cleaned on schedule ongoing

Record what worked and what was noise. Do not extract templates until then.


J. Final simplified standard

Five rules, in this order of importance:

  1. One private Git remote, main plus short-lived branches. Neither the VPS nor WSL is the only copy.
  2. One AGENTS.md per project, containing rules, boundaries and deploy steps. CLAUDE.md is only a pointer. Decisions go into docs/DECISIONS.md.
  3. scripts/check.sh, test.sh light|full, build.sh light|full, plus deploy.sh and rollback.sh for deployed projects. Each script checks the machine role and prints one RESULT: line.
  4. The agent never writes production. Prod is root-owned release directories. The agent drops a checksummed artifact into incoming/, and a human (later a path unit) runs the deploy script.
  5. A host-level disk and RAM ledger. ≥ 2 GB free, ≤ 85% used, one agent session plus one heavy job, agent scratch expires after 3 days.

Everything else — PROJECT_PROFILE.yaml, the bootstrap template, Remote Dev automation of the WSL handoff — waits until two projects have used this and shown where it hurts.


Answers to the 15 questions (short)

  1. Is the VPS/WSL/Windows split sound? Yes. Define it by resource class (light, heavy, native), and enforce it in scripts.
  2. Unnecessarily complicated: PROJECT_PROFILE.yaml, three docs files, a permanent dev branch, pnpm-specific names, per-project budget tables.
  3. Underspecified: the Git remote, the WSL→VPS handoff, version pinning, deploy and migration safety, enforcement of the role split, secrets distribution.
  4. Is Git the right boundary? Yes, with a real remote. Today there is none.
  5. Canonical checkout? Neither. The remote is canonical.
  6. Is AGENTS.md + PROFILE + scripts enough? Nearly. Add decision log, guard in scripts, version pins, and the cold-start drill to prove it.
  7. Minimum commands: check, test light|full, build light|full, and deploy/rollback where deployed.
  8. Heavy-build offloading: WSL initiates; one command produces report and artifact; VPS deploys by artifact id. Automate later.
  9. Prod isolation: low-privilege agent user, root-owned releases, write-only incoming/, human or path-unit deploy, backup before migrate, auto-rollback.
  10. Are the budgets realistic? Per project, yes for one or two Node projects. The host-level floor matters more.
  11. Validate first with Remote Dev: Git remote, light path limits, WSL smoke (Phase 1b), handoff steps, prod isolation.
  12. Remove: see question 2.
  13. Add: remote and backups, role guard, version pins, .gitattributes, secret-scan hook, decision log, scratch cleanup, agent user separation, result channel.
  14. Hidden Claude locks: see H.
  15. Simplify: the five rules in J.