praisonai train remote …) from PraisonAI PR #4367 onward — before that the whole group was unregistered and only worked from a standalone praisonai-train install. It speaks plain ssh/scp through subprocess, so there is no third-party dependency: any host in your ~/.ssh/config works.Quick Start
You have a config, a dataset, and an SSH alias for a GPU box (gpubox in ~/.ssh/config).
Check the host is ready
gpubox is ready to train. — or exits non-zero with the reason (no CUDA GPU, python3 not usable, host unreachable).Start training
Or start and get your shell back
--no-follow if you want the shell back immediately.Check on it later
Fetch the artefacts back
Reattach, inspect, fetch, stop
The run is launched undersetsid (or plain nohup on macOS, which has no setsid), so it outlives the SSH session that started it and its children — dataloader workers, torchrun ranks — sit in one process group stop can take down together. The run id names a directory on the far host; that id is enough to reattach, inspect, fetch, or stop the run later. The launch redirects its stdin from /dev/null, so remote start returns as soon as the trainer is running rather than holding the SSH connection open for the run’s duration.
Everything after start takes only host and run_id.
Deciding which command to run
How It Works
Runs launch undernohup in a detached session. The pid is written to <run_dir>/train.pid and the exit code to <run_dir>/status, so status, tail and stop work from any machine that can ssh to the host.
Commands
Every command takes the SSHhost alias first. Runs are addressed by the run id start prints.
remote preflight
Check the host can train, before shipping anything to it.
remote start
Ship the job and start it. The run outlives this command.
remote tail
Reattach to a running job’s log; prints the final status when it ends.
remote status
Whether a run is still going, and how it ended if not — prints running, completed, failed (exit N), or unknown.
status on a live run could report unknown. The pid file now names the wrapper that owns the trainer, so status probes a live process for the whole run.remote fetch
Bring an artefact back — the adapter, a checkpoint, the log.
remote stop
Stop a run. Signals the whole process group on the host — the wrapper, the trainer, and its dataloader / torchrun children — then re-probes to confirm the group is gone before reporting success. stopped means the run is actually off the GPU, not just that kill -TERM was issued.
stop, reparented to init, and kept the GPU; and stop reported success the moment kill -TERM was issued, letting a UI label a live paid run “cancelled”.Run IDs
--run-id (when passed) must be a plain identifier: letters, digits, dot, dash and underscore only; starting with a letter or digit; up to 64 characters. Anything else is refused:
--run-id was interpolated into the remote path and executed by the far-side shell. Values like x;touch /tmp/pwned, $(id), `id`, a b, ../escape, /absolute, an empty string, and -leading-dash are all refused. run-1787667005 and my_run.2 are fine.
Workdir
--workdir (the remote directory the run lives in, default ~/.praisonai-train) is expanded by the remote shell, so it must survive being passed unquoted. The runner refuses anything outside a safe character set: letters, digits, ., _, -, /, and an optional leading ~. Anything else is rejected before any SSH is opened:
~ is expanded by the remote shell (not the local one), so ~/.praisonai-train lands in the remote user’s real home. The directory is created with mkdir -p -m 700, so on a shared GPU box the umask does not decide who can read the shipped config and dataset.
~ was created in the login directory and every derived path (config, logs, checkpoints) landed there while ~/.praisonai-train stayed empty. If you ran remote training on an older release, log in to the box and check:~ directory (yes, one whose name is a single tilde character) is safe to remove after moving any artefacts you still want out of it:praisonai train remote creates and uses the real ~/.praisonai-train instead.Fetch paths
fetch <run> <path> treats <path> as relative to the run directory on the remote host. These are refused up-front:
- absolute paths (
/etc/passwd) - any path containing
..(../../../../etc/passwd,a/../../b) - empty or whitespace-only paths
shlex.quote()s the scp target, so a shell metacharacter that survives validation cannot be executed on the remote side — validation catches the typo, quoting catches everything else. lora_model and outputs/checkpoint-100 are accepted.
Command shape on the remote side
The remote invocation is:python -m praisonai_train llm config.yaml [DATASET] — that earlier shape produced two positionals, which Typer rejected with exit 2 before the GPU was touched. If a remote start log ends in Error: Got unexpected extra argument on an old version, upgrade to a release carrying PR #4367.
Why “not ready” is not a crash
Preflight failures and stopping something that isn’t running are answers, not errors — but they still exit non-zero (preflight) or printwas not running (stop) so scripts can’t mistake them for success.
Preflight refuses an unready box
Preflight refuses an unready box
Stop is honest about no-ops
Stop is honest about no-ops
Nothing ships without a config
Nothing ships without a config
Prerequisites
Key-based SSH (no password prompts)
Key-based SSH (no password prompts)
ssh in BatchMode=yes — password prompts are never accepted. Set up key auth:~/.ssh/config:A usable Python and GPU on the host
A usable Python and GPU on the host
--python python3 must exist and be able to import praisonai_train. nvidia-smi must list at least --gpus GPUs.Standard OpenSSH only
Standard OpenSSH only
ssh/scp.Python API
Script the same flow withRemoteRunner — each method reconnects, so a runner can be rebuilt from just host + run_id.
Best Practices
Preflight before you ship a big dataset
Preflight before you ship a big dataset
praisonai-train remote preflight gpubox --gpus 2 catches a bad ~/.ssh/config, a missing driver, or a mis-installed Python before you ship a multi-GB dataset.Keep run ids to the safe alphabet
Keep run ids to the safe alphabet
--run-id my_run.2 is fine. Anything a shell would interpret is rejected up-front. Don’t try to be clever — the safe alphabet is intentional.Keep workdirs to the safe alphabet
Keep workdirs to the safe alphabet
--workdir ~/.praisonai-train (the default) is fine. Only letters, digits, ., _, -, /, and an optional leading ~ are allowed; $(pwd), spaces, ;, and quotes are refused before any SSH. See Workdir.Ctrl-C the follow, not the run
Ctrl-C the follow, not the run
start prints the id and stop hint before streaming. If you Ctrl-C the follow, the run keeps going — reattach with remote tail <host> <run_id>.Script polling with status
Script polling with status
start --no-follow returns immediately with the run id. Poll with remote status in a loop, then fetch when it reads completed.Fetch only from the run directory
Fetch only from the run directory
ssh — the CLI won’t do it.Fetch only what you need
Fetch only what you need
fetch uses scp -r — pull lora_model for the adapter or train.log for the log. Avoid pulling the whole run directory.Use a short ~/.ssh/config alias
Use a short ~/.ssh/config alias
Ship the config you want
Ship the config you want
remote start ships the config you pass and overwrites config.yaml on the box too. Don’t rely on the box to already have a matching one.Related
CLI: train
praisonai train subcommands.Train
config.yaml reference for the run that ships to the box.
