Skip to main content
Train on a remote GPU host over plain SSH. The run outlives your SSH session, so closing your laptop does not kill training.
Reachable through the integrated CLI (praisonai train remote …) from PraisonAI PR #4367 onward — before that the whole group was unregistered and only worked from a standalone praisonai-train install. It speaks plain ssh/scp through subprocess, so there is no third-party dependency: any host in your ~/.ssh/config works.

Quick Start

You have a config, a dataset, and an SSH alias for a GPU box (gpubox in ~/.ssh/config).
1

Check the host is ready

Prints gpubox is ready to train. — or exits non-zero with the reason (no CUDA GPU, python3 not usable, host unreachable).
2

Start training

Prints the run id and the stop command before streaming — so a Ctrl-C during the follow still leaves you with the handle:
3

Or start and get your shell back

Use --no-follow if you want the shell back immediately.
4

Check on it later

Every command reattaches from just the host and run id, so a dropped laptop connection never kills training.
5

Fetch the artefacts back

Pull the adapter (or any file inside the run directory) home.

Reattach, inspect, fetch, stop

The run is launched under setsid (or plain nohup on macOS, which has no setsid), so it outlives the SSH session that started it and its children — dataloader workers, torchrun ranks — sit in one process group stop can take down together. The run id names a directory on the far host; that id is enough to reattach, inspect, fetch, or stop the run later. The launch redirects its stdin from /dev/null, so remote start returns as soon as the trainer is running rather than holding the SSH connection open for the run’s duration. Everything after start takes only host and run_id.

Deciding which command to run


How It Works

Runs launch under nohup in a detached session. The pid is written to <run_dir>/train.pid and the exit code to <run_dir>/status, so status, tail and stop work from any machine that can ssh to the host.

Commands

Every command takes the SSH host alias first. Runs are addressed by the run id start prints.

remote preflight

Check the host can train, before shipping anything to it.

remote start

Ship the job and start it. The run outlives this command.

remote tail

Reattach to a running job’s log; prints the final status when it ends.

remote status

Whether a run is still going, and how it ended if not — prints running, completed, failed (exit N), or unknown.
Fixed in PraisonAI PR #4553 (closes #4549). Earlier releases recorded the pid of a transient wrapper subshell that could exit while training kept running, so status on a live run could report unknown. The pid file now names the wrapper that owns the trainer, so status probes a live process for the whole run.

remote fetch

Bring an artefact back — the adapter, a checkpoint, the log.

remote stop

Stop a run. Signals the whole process group on the host — the wrapper, the trainer, and its dataloader / torchrun children — then re-probes to confirm the group is gone before reporting success. stopped means the run is actually off the GPU, not just that kill -TERM was issued.
Three outcomes:
Fixed in PraisonAI PR #4551 (closes #4547) — the remote counterpart of #4491. Earlier releases signalled only the recorded pid (the wrapper shell), so real fine-tune children survived stop, reparented to init, and kept the GPU; and stop reported success the moment kill -TERM was issued, letting a UI label a live paid run “cancelled”.
The recording side of the same class of bug is fixed in PraisonAI PR #4553 (closes #4549): the pid file now names the wrapper whose group contains the trainer, so the group stop resolves is the one the trainer actually runs in.

Run IDs

--run-id (when passed) must be a plain identifier: letters, digits, dot, dash and underscore only; starting with a letter or digit; up to 64 characters. Anything else is refused:
This validator was added in PR #4367: a shell metacharacter inside --run-id was interpolated into the remote path and executed by the far-side shell. Values like x;touch /tmp/pwned, $(id), `id`, a b, ../escape, /absolute, an empty string, and -leading-dash are all refused. run-1787667005 and my_run.2 are fine.

Workdir

--workdir (the remote directory the run lives in, default ~/.praisonai-train) is expanded by the remote shell, so it must survive being passed unquoted. The runner refuses anything outside a safe character set: letters, digits, ., _, -, /, and an optional leading ~. Anything else is rejected before any SSH is opened:
The ~ is expanded by the remote shell (not the local one), so ~/.praisonai-train lands in the remote user’s real home. The directory is created with mkdir -p -m 700, so on a shared GPU box the umask does not decide who can read the shipped config and dataset.
Migration note (PraisonAI PR #4550). Before this fix, the workdir was quoted before the remote shell saw it, so a directory literally named ~ was created in the login directory and every derived path (config, logs, checkpoints) landed there while ~/.praisonai-train stayed empty. If you ran remote training on an older release, log in to the box and check:
A ~ directory (yes, one whose name is a single tilde character) is safe to remove after moving any artefacts you still want out of it:
On the current release, praisonai train remote creates and uses the real ~/.praisonai-train instead.

Fetch paths

fetch <run> <path> treats <path> as relative to the run directory on the remote host. These are refused up-front:
  • absolute paths (/etc/passwd)
  • any path containing .. (../../../../etc/passwd, a/../../b)
  • empty or whitespace-only paths
The runner also shlex.quote()s the scp target, so a shell metacharacter that survives validation cannot be executed on the remote side — validation catches the typo, quoting catches everything else. lora_model and outputs/checkpoint-100 are accepted.

Command shape on the remote side

The remote invocation is:
Not python -m praisonai_train llm config.yaml [DATASET] — that earlier shape produced two positionals, which Typer rejected with exit 2 before the GPU was touched. If a remote start log ends in Error: Got unexpected extra argument on an old version, upgrade to a release carrying PR #4367.

Why “not ready” is not a crash

Preflight failures and stopping something that isn’t running are answers, not errors — but they still exit non-zero (preflight) or print was not running (stop) so scripts can’t mistake them for success.
A script reading only the exit status will not ship a job to a box that can’t run it.
Reporting success for a no-op would let a script believe it killed a run still burning GPU hours.
The missing config is refused before any SSH happens.

Prerequisites

The runner uses ssh in BatchMode=yes — password prompts are never accepted. Set up key auth:
Alias config lives in ~/.ssh/config:
The default --python python3 must exist and be able to import praisonai_train. nvidia-smi must list at least --gpus GPUs.
No paramiko, fabric, or SSH library required — the runner shells out to ssh/scp.

Python API

Script the same flow with RemoteRunner — each method reconnects, so a runner can be rebuilt from just host + run_id.

Best Practices

praisonai-train remote preflight gpubox --gpus 2 catches a bad ~/.ssh/config, a missing driver, or a mis-installed Python before you ship a multi-GB dataset.
--run-id my_run.2 is fine. Anything a shell would interpret is rejected up-front. Don’t try to be clever — the safe alphabet is intentional.
--workdir ~/.praisonai-train (the default) is fine. Only letters, digits, ., _, -, /, and an optional leading ~ are allowed; $(pwd), spaces, ;, and quotes are refused before any SSH. See Workdir.
start prints the id and stop hint before streaming. If you Ctrl-C the follow, the run keeps going — reattach with remote tail <host> <run_id>.
start --no-follow returns immediately with the run id. Poll with remote status in a loop, then fetch when it reads completed.
Fetch is scoped to the run’s own artefacts. If you need something outside the run directory, log into the box directly with ssh — the CLI won’t do it.
fetch uses scp -r — pull lora_model for the adapter or train.log for the log. Avoid pulling the whole run directory.
Everything the CLI does is one hop, per-alias. A short entry with the host, user, and identity keeps the CLI arguments short.
remote start ships the config you pass and overwrites config.yaml on the box too. Don’t rely on the box to already have a matching one.

CLI: train

All praisonai train subcommands.

Train

config.yaml reference for the run that ships to the box.

Train Export

Publish the adapter you fetched back — to Hugging Face, GGUF, or Ollama.

Multi-GPU Training

Fine-tune across every GPU on the host with torchrun.

Checkpointing

Resume from checkpoints if a long run dies mid-way.