Hybrid AI Development Model — Claude Review
Plan review only. Nothing was created, installed, cleaned, deployed or configured. The only file written is this review, in the session scratchpad. Read-only checks were run on the VPS to test the model against what already exists.
What I observed on the VPS (read-only)
| Observation | Why it matters |
|---|---|
The Lucky.ai dev repo /home/claude/dev/luckyai is a git repo (68 commits, branch ci/case-intelligence) with no remote configured |
The premise "private Git repo is the source of truth" is not true yet. The VPS disk is currently the only copy |
Prod /home/claude/luckyai is not a git repo and is owned by claude (755, plus env files with .pre-* backups beside them) |
The agent user can edit live production. This is exactly the failure the plan says to avoid |
/home/claude/case, luckyai-data, cdesktop-runtime are not git repos |
Unversioned data and code exist next to versioned code |
No sudo, no gh, no git credential helper, one SSH private key present in ~/.ssh (I did not check where it points) |
The agent cannot deploy as root. A remote git host is not set up yet |
Lucky.ai dev has README.md, docs/, scripts/ but no AGENTS.md or CLAUDE.md. cdesktop-custom has both files, side by side |
Agent rules are missing in one project and duplicated in another |
Existing scripts: luckyai-dev.sh, backup-now.sh, restore-vps.sh, verify-backup.sh, ci-*.sh. No check/test/build/deploy names, no deploy script |
A good base to standardise from, but no contract yet |
~/.claude/settings.json holds model, status line, plugins. Skills live in ~/.claude/skills and ~/.agent/skills. Claude memory holds user preferences |
Some working knowledge lives outside Git, in Claude-specific places |
| Host is 2 GB RAM, 1 vCPU, about 2.0 GB free disk (90% used). Each Claude process is 300–400 MB RSS | Budgets must be host-level, not only per project |
| Lucky.ai already runs Playwright on the VPS (one browser cache is 300–650 MB) | The plan's "Playwright → WSL" rule conflicts with current practice |
Earlier today /tmp/claude-1000 held about 1.5 GB of agent scratch. 705 MB of it was two Python venvs |
Agent leftovers, not projects, were the biggest disk cost I found |
A. Verdict
PASS WITH CHANGES
The core idea is sound: one Git remote, a light VPS, a heavy WSL, a native Windows, and the agent treated as replaceable. Four things need fixing before it can be called a standard:
- Git has no remote yet.
- The agent can still write production.
- The "where does this run" policy is advisory text, not enforced.
- Too much project knowledge still lives in Claude-specific places.
B. Key strengths
- Git as the boundary, with separate dev and prod checkouts, is right and cheap.
- "Move heavy work elsewhere, don't upgrade the VPS" matches the measured constraints (RAM and disk are the binding limits).
- Agent-as-replaceable-developer is the correct framing. Putting knowledge into repo files, not chat history, is the right lever.
- Early WSL smoke tests avoid late surprises. The Remote Dev plan already has this as Phase 1b.
- Deferring the bootstrap template until two projects have proven the workflow is the right order.
- Release directories with rollback fit the small-VPS reality.
C. Problems / risks
High
- No Git remote, so no source of truth. The VPS is meant to be "replaceable", yet today it holds the only copy of the dev repo. Fix: create a private remote first, and push from both VPS and WSL. Hosted private repo is the cheapest choice. A self-hosted Gitea would cost RAM, and a bare repo on WSL is unreachable while WSL sleeps.
- The agent can write production. Prod is owned by the same user the agent runs as, and there is no
sudo. The plan says "AI agents should not casually edit production files", but nothing enforces it. Fix: agent runs as a low-privilege user; prod is owned by root (or adeployuser); the agent can only drop an artifact into a write-onlyincoming/directory. Deploy is triggered by a human at first, and later by a systemd path unit that runs a root-owned deploy script. That needs nosudo. - The VPS/WSL/Windows policy is documentation, not enforcement. "The AI agent should not need to guess" is only true if the scripts refuse to run in the wrong place. Fix: each machine declares its role (
DEV_ENV=vps|wsl|windows, one line in a profile file), andscripts/share a small guard that refuses heavy work on the VPS unless explicitly overridden. The guard lives in the repo, so it is agent-independent. - The WSL → VPS handoff is undefined. Step 7 ("problems are reported back") and the artifact flow have no channel. WSL sleeps and sits behind NAT, so it must always initiate. See D for a minimal design.
Medium
pnpmcommand names assume Node projects. Lucky.ai is Python (FastAPI) plus React. Usescripts/*.shas the contract; they can call pnpm, pytest or anything else.- Reproducibility across machines is unspecified. "Works on WSL, fails on VPS" is a common trap. Pin Node and Python versions (
.nvmrc,.python-version), commit lockfiles, install with--frozen-lockfile. PROJECT_PROFILE.yamlduplicatesAGENTS.mduntil something reads it. Two files that say the same thing drift.- Deploys and database migrations. A rollback that swaps a symlink does not undo a forward migration. Deploy must take a backup first (Lucky.ai already has
backup-now.shandrestore-vps.sh), and migrations must be backward-compatible or gated. - Per-project disk budgets miss the real consumers: package caches, Python venvs (Lucky.ai has three), Playwright browsers, agent scratch, session transcripts (about 229 MB under
~/.claude), backups. See G. - Hidden Claude couplings remain (see H).
Low
- Line endings: Windows checkouts can turn shell scripts into CRLF. Add
.gitattributeswith* text=auto eol=lf. - Android builds do not strictly need Windows. Gradle can run in WSL with the command-line SDK. Windows is needed for the emulator and GUI tools. Keep Windows as an opt-in role for native projects only.
devas a permanent branch adds merge work for a one-person flow.mainplus short-lived branches is enough.
D. Recommended architecture (corrected model)
private Git remote (canonical: nobody's disk is)
▲ ▲ ▲
push/pull│ │ │push/pull
│ │ │
VPS ─────────┘ │ └───────── WSL
agent (low-priv user) │ heavy build, full tests,
edit, check, light test │ Codex, browser tests
light build │ ── verify report ──► VPS incoming/
│ │ ── artifact ───────► VPS incoming/
▼ │
incoming/ (write-only) │ Windows: only when a project needs
│ native tools (Android emulator, GUI).
▼ (human first; systemd path unit later) Same Git remote, own checkout.
deploy.sh (root-owned)
verify checksum → backup → unpack → migrate → health check → swap → verify → rollback on failure
▼
/opt/<project>/releases/<id> + current (root-owned; agent cannot write)
Handoff design (minimum, no automation platform)
- WSL → VPS is always initiated by WSL. WSL is behind NAT and sleeps.
- One command on WSL:
scripts/wsl-verify.sh [commit]. It fetches the commit, runs full test and full build, writesreport-<commit>.mdand, if requested,<project>-<commit>.tgzplus a checksum. It then copies both to the VPSincoming/over SSH (outbound only). - The VPS agent reads the report from
incoming/. Until Remote Dev automates this, the human runs one command and the agent reads one file. - Deploy takes an artifact id. It verifies the checksum and that the commit exists on
main, then follows the pipeline above. - Do not automate further until the manual flow has shown real pain.
Which checkout is canonical
The Git remote is canonical. Neither checkout is. The VPS checkout is the agent's working copy, and WSL is a verifier that does not commit to shared branches. WSL-only fixes go on wsl/* branches, as the plan says. This also makes the failure model work: if the VPS agent is unavailable, WSL can keep developing against the same remote.
E. VPS / WSL / Windows responsibility matrix
| Task | VPS | WSL | Windows |
|---|---|---|---|
| Agent editing, git commit/push | ✅ primary | fallback | — |
check (lint, typecheck) |
✅ | ✅ | — |
| Light tests (unit, small integration) | ✅ | ✅ | — |
| Light build | ✅ (memory-capped) | ✅ | — |
| Full tests, browser tests, stress tests | ❌ | ✅ | — |
| Full build, large installs | ❌ | ✅ | — |
| Codex, real local runner tests | ❌ | ✅ | — |
| Production runtime, deploy | ✅ (root-owned) | ❌ | ❌ |
| Android emulator, Android Studio, GUI tools | ❌ | ❌ | ✅ |
| APK/AAB build | ❌ | optional (CLI SDK) | ✅ default |
| Secrets | per-machine .env |
per-machine .env |
per-machine |
Exception to record: Lucky.ai already runs headless Playwright on the VPS. Either move it to WSL, or allow "headless shell only, one at a time, memory-capped" as a written exception in that project's AGENTS.md.
F. Recommended standard project contract
Required (every project):
README.md what it is, how to run it
AGENTS.md the single source of agent rules (see below)
scripts/check.sh fast, VPS-safe: lint + typecheck (target < 60 s)
scripts/test.sh takes light|full
scripts/build.sh takes light|full
.env.example variable names only
.gitattributes eol=lf
version pins + lockfile .nvmrc / .python-version, pnpm-lock / requirements lock
Required only for projects that deploy to the VPS:
scripts/deploy.sh artifact id in; checksum, backup, unpack, health check, swap, rollback
scripts/rollback.sh
docs/ARCHITECTURE.md one file is enough
Optional: Makefile as sugar (make is already installed); PROJECT_PROFILE.yaml once a tool consumes it.
Script conventions:
- Each script runs the machine-role guard first (
DEV_ENV), and a heavy target on the VPS refuses unlessFORCE_HEAVY=1. - The last output line is machine-readable:
RESULT: PASS|FAIL step=<name> where=<vps|wsl> cost=<light|heavy>. Remote Dev can consume it later.
AGENTS.md must contain: purpose; repo and directory boundaries; what runs where (points to the guard, not a second copy of it); test and deploy procedure; forbidden actions (destructive commands, editing prod, committing secrets); operating rules (ask before deleting or publishing; long outputs go to a file); a pointer to docs/DECISIONS.md.
CLAUDE.md, where needed, should be a one-line pointer to AGENTS.md (or import it, as I understand Claude Code allows). Keep one copy. The cdesktop-custom project today keeps two copies that can drift.
G. Resource policy (host-level first)
The binding limit is the host, not any one project. At about 2.0 GB free, a single 700 MB project already uses a third of the headroom.
Host rules
| Rule | Value |
|---|---|
| Disk floor | keep ≥ 2 GB free and ≤ 85% used. Today's 90% is already over the line |
| Alert | 90% used, or under 1.5 GB free |
| Disk ledger | monthly du per project and per agent area, one line each, in a short file |
| Agent scratch | forbid venvs and node_modules in scratch; agent temp older than 3 days is disposable (this alone reclaimed 705 MB this week) |
| Package caches | trim monthly (~/.npm about 98 MB now) |
| Browsers | one Playwright install per host, never deleted while any session runs |
| RAM | at most 1 interactive agent session + 1 heavy job; builds run with NODE_OPTIONS=--max-old-space-size=768 and nice; consider systemd-run --scope -p MemoryMax= |
| Swap | treat sustained swap use (46% today) as a warning sign, not free memory |
Per-project budgets (keep, but they were realistic only for Node projects):
| Item | Default | Add |
|---|---|---|
| Source + git | < 50 MB | |
| Dev dependencies | < 250 MB | Python venvs count too (Lucky.ai has three) |
| Production dependencies | < 100 MB | keep the last 2 releases only |
| Build output | < 50 MB | |
| Logs | < 50 MB | journald limit |
| Browsers | not in project budget | host-level, one copy |
| Backups | not in project budget | host-level, with a retention count |
The 300–700 MB per project target is realistic for one or two small projects. It is not a safe multiplier for many projects on this host.
H. Agent-replaceability review
Verdict: the infrastructure can already survive replacing Claude. The knowledge cannot yet.
Infrastructure (git, scripts, deploy, systemd, Caddy) has nothing Claude-specific. What still locks the project to Claude:
| Hidden dependency | Where it lives | Fix |
|---|---|---|
| Project decisions and reasoning | Claude session transcripts (about 229 MB under ~/.claude) |
Keep docs/DECISIONS.md: one short entry per non-obvious decision |
| Working preferences (ask before acting, long output as files) | Claude memory | Mirror project-relevant ones in AGENTS.md; keep the rest as personal preferences |
| Skills and plugins | ~/.claude/skills, ~/.agent/skills, settings.json |
Treat as tooling. Anything a project needs must be a script in the repo |
| Duplicate agent files | AGENTS.md and CLAUDE.md both hand-written |
One real file, one pointer |
| Model IDs and effort settings | env files (.pre-opus55, .pre-sonnet55 copies) and models.conf |
Keep them in config with a comment, not in habit |
| Permission behaviour (sessions run with dangerous-mode flags) | Claude CLI semantics | Enforce with OS user, script guards and root-owned prod, which work for any agent |
| Lucky.ai's product runtime uses Claude Code | runner service | This is a product dependency, not a development one. Track separately, behind the adapter that the Remote Dev plan already proposes |
Test it, don't assume it: run a cold-start drill. Give a fresh agent (Codex, or Claude with no memory) one small task and only the repository. If it cannot finish using README.md, AGENTS.md and the scripts, the standard is not yet agent-independent. Codex reads AGENTS.md natively, as I understand it. Verify that when you first try it.
I. Remote Dev pilot plan
Use the Remote Dev milestones as the pilot. Each hypothesis has a pass/fail check.
| Hypothesis | Check | When |
|---|---|---|
| H1: the light path fits the VPS | check under 60 s, test light and build light stay under the 768 MB cap, host free disk does not drop below the floor |
M1 |
| H2: Git is a sufficient boundary | Private remote exists; VPS pushes, WSL pulls the same commit, and a wsl/* fix merges back cleanly |
M1 (needs the remote first) |
| H3: WSL catches what VPS cannot | Phase 1b smoke test: at least one real difference is found and recorded (paths, sleep, CLI behaviour) | Phase 1b |
| H4: the handoff is one command each side | wsl-verify.sh on WSL, then the agent reads report-<commit>.md. Count the manual steps; the goal is 1 human action |
Phase 1b–2 |
| H5: prod is isolated from the agent | Agent user cannot write /opt/remote-dev. Deploy succeeds through incoming/, a forced failure rolls back |
Phase 5 |
| H6: an agent can be replaced | Cold-start drill with a second agent, repo only | after M1 |
| H7: resource policy holds | Monthly ledger: disk, RAM peaks, scratch cleaned on schedule | ongoing |
Record what worked and what was noise. Do not extract templates until then.
J. Final simplified standard
Five rules, in this order of importance:
- One private Git remote,
mainplus short-lived branches. Neither the VPS nor WSL is the only copy. - One
AGENTS.mdper project, containing rules, boundaries and deploy steps.CLAUDE.mdis only a pointer. Decisions go intodocs/DECISIONS.md. scripts/check.sh,test.sh light|full,build.sh light|full, plusdeploy.shandrollback.shfor deployed projects. Each script checks the machine role and prints oneRESULT:line.- The agent never writes production. Prod is root-owned release directories. The agent drops a checksummed artifact into
incoming/, and a human (later a path unit) runs the deploy script. - A host-level disk and RAM ledger. ≥ 2 GB free, ≤ 85% used, one agent session plus one heavy job, agent scratch expires after 3 days.
Everything else — PROJECT_PROFILE.yaml, the bootstrap template, Remote Dev automation of the WSL handoff — waits until two projects have used this and shown where it hurts.
Answers to the 15 questions (short)
- Is the VPS/WSL/Windows split sound? Yes. Define it by resource class (light, heavy, native), and enforce it in scripts.
- Unnecessarily complicated:
PROJECT_PROFILE.yaml, three docs files, a permanentdevbranch, pnpm-specific names, per-project budget tables. - Underspecified: the Git remote, the WSL→VPS handoff, version pinning, deploy and migration safety, enforcement of the role split, secrets distribution.
- Is Git the right boundary? Yes, with a real remote. Today there is none.
- Canonical checkout? Neither. The remote is canonical.
- Is
AGENTS.md + PROFILE + scriptsenough? Nearly. Add decision log, guard in scripts, version pins, and the cold-start drill to prove it. - Minimum commands:
check,test light|full,build light|full, anddeploy/rollbackwhere deployed. - Heavy-build offloading: WSL initiates; one command produces report and artifact; VPS deploys by artifact id. Automate later.
- Prod isolation: low-privilege agent user, root-owned releases, write-only
incoming/, human or path-unit deploy, backup before migrate, auto-rollback. - Are the budgets realistic? Per project, yes for one or two Node projects. The host-level floor matters more.
- Validate first with Remote Dev: Git remote, light path limits, WSL smoke (Phase 1b), handoff steps, prod isolation.
- Remove: see question 2.
- Add: remote and backups, role guard, version pins,
.gitattributes, secret-scan hook, decision log, scratch cleanup, agent user separation, result channel. - Hidden Claude locks: see H.
- Simplify: the five rules in J.