Hillclimbing
Agentic model searches — agents draft, backtest, and select models; you promote and deploy the winner.
rebase hillclimb runs agentic model searches with
hillclimb: coding
agents draft, debug, improve, and ensemble
emflow Predictor classes. Every
candidate is backtested leakage-safe on the problem's validation split, and the
shipped model is selected on a hidden holdout the agents never see.
The artifact that flows through the whole loop is a single Python file
exposing get_model() -> emflow.Predictor — the same class that won the
backtest is what you commit to your workspace repo and deploy.
Requires the hillclimb extra:
uv pip install "rebase-toolkit[hillclimb]"Start a search
Hosted on the platform by default (a long-running Cloud Run job), or on your
own machine with --local:
rebase hillclimb start emflow://gefcom2014:solar --budget 2h
rebase hillclimb start emflow://gefcom2014:solar --budget 2h --parallel-searches 3
rebase hillclimb init # once per repo, before local searches: hillclimb/config.yaml + hillclimb/runs/
rebase hillclimb start emflow://gefcom2014:solar --budget 2h --localHosted searches take emflow registry problems (emflow://<name>); list them with
rebase hillclimb problems. Plain problem folders
run with --local. Useful options:
| Option | Description |
|---|---|
--budget | Wall-clock search budget, e.g. 2h, 30m. |
--name | Search name shown in run listings. |
--model | Agent model, e.g. sonnet. |
--policy | Search engine: greedy (default), openevolve, gepa. |
--parallel-searches | Independent searches under one hosted run, sharing knowledge as they go. |
--parallel-operators | Agents each search keeps busy at once (hosted default 3). The job is sized from searches × operators. |
--claude-secret | Workspace secret that bills the agents (default: hillclimb when it exists). |
--backend dummy | Run the search loop without agent calls (smoke tests). |
--holdout / --no-holdout | Score candidates on the holdout split for selection (default on). |
--local | Run on this machine instead of the platform. Honors --budget, --name, --model, --backend and --no-holdout; policy, parallelism and secrets come from hillclimb/config.yaml. |
--project | Project the platform run is filed under (default hillclimb). |
Watch and control
A hosted search mirrors its state (status heartbeats, the candidate journal,
agent streams, best/) to the platform every few seconds. The engine's own
TUIs open on a live copy of it — the same screens as the standalone
hillclimb watch, and the stop and prune keys reach the running job:
rebase hillclimb watch <run-id> # runs → searches → candidates; s = stop, x = prune
rebase hillclimb chart <run-id> # best score so far, every candidate a dot
rebase hillclimb tree <run-id> # one search's exploration tree
rebase hillclimb graph <run-id> # the knowledge graph
rebase hillclimb status <run-id> # plain text: candidates, best score, budget left
rebase hillclimb logs <run-id> # engine logs (one per search)
rebase hillclimb stop <run-id> # graceful: parks after the current operator
rebase hillclimb listWithout a run id the TUI commands open the local hillclimb dir. rebase run get <run-id> shows the underlying platform run and its events.
How a search works
- A baseline candidate (the problem's reference model, when it ships one) is evaluated first — agents must beat a real scored floor.
- Agents draft diverse solutions, debug failures, and improve the best candidate; near the end of the budget an ensemble operator combines the top candidates.
- Every candidate is fit and scored by a generic evaluator on the validation split — sandboxed, offline, and credential-free, so generated code can never read holdout data or secrets.
- The orchestrator separately scores each ok candidate on the holdout split and selects the shipped model by a rank-blend of both signals — an overfit validation score gets vetoed by its holdout rank.
- A finished search runs one official emflow verification on the selected
model, recording the number of candidates tried (
n_trials) for selection honesty.
Promote and deploy
Fetch the selected model into your workspace repo as versioned source, then deploy it like any other model:
rebase hillclimb promote <run-id> # writes models/<search>.py per search, e.g. gefcom2014_solar.py; --pr opens the PR
# review, commit, open a PR (protected environments deploy through gitops)
rebase model deploy models/gefcom2014_solar.pyBilling and state
Hosted agents bill the workspace secret named hillclimb when it exists — a
CLAUDE_CODE_OAUTH_TOKEN for subscription billing, or an ANTHROPIC_API_KEY —
and otherwise the platform's own credentials. Create it once per workspace:
claude setup-token | rebase secret create hillclimb CLAUDE_CODE_OAUTH_TOKEN=---claude-secret names a different bundle for one search. --local searches
use your local Claude login. Credits cover the job's compute for the requested
budget plus a per-search agent-spend ceiling; the engine parks the search
(resumable) when it reaches the ceiling.
Search state is served by the platform (GET /runs/{id}/hillclimb/objects,
POST /runs/{id}/hillclimb/control); nothing on your machine needs Google
credentials or a bucket name. Server-side requirements (job image, secrets,
artifacts bucket) are documented in the platform operations guide
(toolkit/HILLCLIMB.md).

