praisonai-train llm job on another machine you reach over SSH — the tiny local CLI stays on your laptop, the GPU box does the heavy work, and the artefacts come back to you.
Reachable through the integrated CLI (
praisonai train remote …) from PraisonAI PR #4367 onward — before that the whole group was unregistered and only worked from a standalone praisonai-train install. It speaks plain ssh/scp through subprocess, so there is no third-party dependency: any host in your ~/.ssh/config works.Quick Start
1
Check the host and start a run
Have SSH access to a machine with a GPU and a working
praisonai-train[llm] install. Point the CLI at it by its SSH alias. start prints a run id and streams the log until the run ends.2
Check on it later
Every command reattaches from just the host and run id, so a dropped laptop connection never kills training.
3
Fetch the artefacts back
Pull the adapter (or any file inside the run directory) home.
How it works
The run is launched undersetsid (or plain nohup on macOS, which has no setsid), so it outlives the SSH session that started it and its children — dataloader workers, torchrun ranks — sit in one process group stop can take down together. The run id names a directory on the far host; that id is enough to reattach, inspect, fetch, or stop the run later. The launch redirects its stdin from /dev/null, so remote start returns as soon as the trainer is running rather than holding the SSH connection open for the run’s duration.
Before shipping anything, start runs a preflight: the host must be reachable, its Python importable, and nvidia-smi must report at least --gpus GPUs. A shortfall refuses the run rather than wasting an upload.
Commands
Every command takes the SSHhost alias first. Runs are addressed by the run id start prints.
remote preflight
Check the host can train, before shipping anything to it.
remote start
Ship the job and start it. The run outlives this command.
remote tail
Reattach to a running job’s log; prints the final status when it ends.
remote status
Whether a run is still going, and how it ended if not — prints running, completed, failed (exit N), or unknown.
Fixed in PraisonAI PR #4553 (closes #4549). Earlier releases recorded the pid of a transient wrapper subshell that could exit while training kept running, so
status on a live run could report unknown. The pid file now names the wrapper that owns the trainer, so status probes a live process for the whole run.remote fetch
Bring an artefact back — the adapter, a checkpoint, the log.
remote stop
Stop a run. Signals the whole process group on the host — the wrapper, the trainer, and its dataloader / torchrun children — then re-probes to confirm the group is gone before reporting success. stopped means the run is actually off the GPU, not just that kill -TERM was issued.
Fixed in PraisonAI PR #4551 (closes #4547) — the remote counterpart of #4491. Earlier releases signalled only the recorded pid (the wrapper shell), so real fine-tune children survived
stop, reparented to init, and kept the GPU; and stop reported success the moment kill -TERM was issued, letting a UI label a live paid run “cancelled”.Run IDs
--run-id (when passed) must be a plain identifier: letters, digits, dot, dash and underscore only; starting with a letter or digit; up to 64 characters. Anything else is refused:
--run-id was interpolated into the remote path and executed by the far-side shell. Values like x;touch /tmp/pwned, $(id), `id`, a b, ../escape, /absolute, an empty string, and -leading-dash are all refused. run-1787667005 and my_run.2 are fine.
Workdir
--workdir (the remote directory the run lives in, default ~/.praisonai-train) is expanded by the remote shell, so it must survive being passed unquoted. The runner refuses anything outside a safe character set: letters, digits, ., _, -, /, and an optional leading ~. Anything else is rejected before any SSH is opened:
~ is expanded by the remote shell (not the local one), so ~/.praisonai-train lands in the remote user’s real home. The directory is created with mkdir -p -m 700, so on a shared GPU box the umask does not decide who can read the shipped config and dataset.
Migration note (PraisonAI PR #4550). Before this fix, the workdir was quoted before the remote shell saw it, so a directory literally named A On the current release,
~ was created in the login directory and every derived path (config, logs, checkpoints) landed there while ~/.praisonai-train stayed empty. If you ran remote training on an older release, log in to the box and check:~ directory (yes, one whose name is a single tilde character) is safe to remove after moving any artefacts you still want out of it:praisonai train remote creates and uses the real ~/.praisonai-train instead.Fetch paths
fetch <run> <path> treats <path> as relative to the run directory on the remote host. These are refused up-front:
- absolute paths (
/etc/passwd) - any path containing
..(../../../../etc/passwd,a/../../b) - empty or whitespace-only paths
shlex.quote()s the scp target, so a shell metacharacter that survives validation cannot be executed on the remote side — validation catches the typo, quoting catches everything else. lora_model and outputs/checkpoint-100 are accepted.
Command shape on the remote side
The remote invocation is:python -m praisonai_train llm config.yaml [DATASET] — that earlier shape produced two positionals, which Typer rejected with exit 2 before the GPU was touched. If a remote start log ends in Error: Got unexpected extra argument on an old version, upgrade to a release carrying PR #4367.
Best Practices
Keep run ids to the safe alphabet
Keep run ids to the safe alphabet
--run-id my_run.2 is fine. Anything a shell would interpret is rejected up-front. Don’t try to be clever — the safe alphabet is intentional.Keep workdirs to the safe alphabet
Keep workdirs to the safe alphabet
--workdir ~/.praisonai-train (the default) is fine. Only letters, digits, ., _, -, /, and an optional leading ~ are allowed; $(pwd), spaces, ;, and quotes are refused before any SSH. See Workdir.Fetch only from the run directory
Fetch only from the run directory
Fetch is scoped to the run’s own artefacts. If you need something outside the run directory, log into the box directly with
ssh — the CLI won’t do it.Use a short ~/.ssh/config alias
Use a short ~/.ssh/config alias
Everything the CLI does is one hop, per-alias. A short entry with the host, user, and identity keeps the CLI arguments short.
Ship the config you want
Ship the config you want
remote start ships the config you pass and overwrites config.yaml on the box too. Don’t rely on the box to already have a matching one.Related
CLI: train
All
praisonai train subcommands.Train
config.yaml reference for the run that ships to the box.
