Buckets

Object storage addressed by key with rb.Bucket, backed one-to-one by a cloud bucket.

A bucket is object storage: keys in, objects out. It is the default place to put data that outlives a run — parquet files, exports, model artifacts, anything you would otherwise reach for S3 for.

import rebase as rb

b = rb.Bucket.from_name("forecasts", create_if_missing=True)

b.put("2026/08/10.parquet", data)
b.get("2026/08/10.parquet")

for obj in b.iter_all(prefix="2026/"):
    print(obj.key, obj.size)

Each rb.Bucket maps to exactly one real cloud bucket. That is what keeps bucket-level behaviour available to you later — lifecycle rules, versioning, per-bucket location — none of which can be expressed on a shared bucket carved up by prefix.

Bucket or Volume?

Both store bytes in the same cloud object storage. They differ in what they pretend to be.

rb.Bucketrb.Volume
ShapeKeys and objectsA mounted directory
Access in a rungs:// paths, or SDK callsOrdinary open() on a path
Writing one byte of a large objectOne requestDownloads and re-uploads the whole object
RenameCopy then delete, explicitlyLooks atomic, is not
File locking, appends, partial writesNot offeredAppear to work, mostly don't
Good forData files, exports, artifacts, anything written onceRead-mostly artifacts where a library demands a real path

Reach for a Bucket by default. Reach for a Volume only when code you do not control insists on a filesystem path. A volume is a gcsfuse mount, and a mount makes expensive operations look cheap: a one-byte edit to a 1 GB file transfers 2 GB, and nothing in the API tells you so.

Attaching a Bucket to Deployed Code

@rb.function(project="energy", mode="job", buckets=["forecasts"])
def train(site_id: str = "site-001") -> dict:
    import pandas as pd

    uri = rb.Bucket.from_name("forecasts").uri
    df = pd.read_parquet(f"{uri}/2026/08/10.parquet")
    return {"rows": len(df)}

Attaching injects REBASE_BUCKET_FORECASTS into the container, holding the bucket's gs:// URI, and grants the runtime service account access to it. Inside a run, gs:// paths work directly through ambient credentials — no signed URLs, no round trip through the API, full throughput. Bucket.uri reads that variable, so it costs nothing.

The environment variable is derived from the bucket name: uppercased, with anything that isn't a letter or digit becoming _. So raw-data.v2 becomes REBASE_BUCKET_RAW_DATA_V2. Two attached buckets that collide on that name are rejected at deploy time rather than silently overwriting each other.

Buckets attach to functions and ASGI apps. A function with buckets must use isolation="dedicated" or mode="job"; the shared interactive runner receives no attachment environment, so that combination is rejected at deploy time.

Ephemeral runs carry most, not all, attachments. rebase run ./file.py::fn creates an ephemeral run from local source. It carries the function's env, secrets and buckets= grants, so it may read and write its attached buckets — but it gets no volumes.

Don't depend on os.environ["REBASE_BUCKET_…"] in code you also run ephemerally; use rb.Bucket.from_name("raw-data").uri instead. Inside a run it reads the injected variable when present and otherwise looks the URI up once, so the same code works deployed, ephemeral, and on your own machine.

From Your Machine

The SDK moves bytes over presigned URLs — uploads and downloads go directly between you and object storage, not through the platform API:

b = rb.Bucket.from_name("forecasts")

b.put_file("./model.pkl", "models/model.pkl")
b.put_directory("./exports", prefix="exports/")
b.download("models/model.pkl", "./model.pkl")

Metadata and deletion go through the API, which is also where permissions are enforced:

b.stat("models/model.pkl")
b.exists("models/model.pkl")
b.delete("models/model.pkl")   # deletes one OBJECT — the bucket itself is deleted
                               # with `rebase bucket delete` or Client.delete_bucket()

Listing is paginated, and says so. list() returns one page plus a next_page_token, so you can always tell a full page from a complete listing. iter_all() follows the pages for you:

page = b.list(prefix="2026/", delimiter="/")
page["objects"]           # objects directly under 2026/
page["prefixes"]          # "folders": 2026/08/, 2026/09/, ...
page["next_page_token"]   # None when this was the last page

for obj in b.iter_all(prefix="2026/"):
    ...

A directory upload signs its URLs in one batch rather than one request per file, so uploading a few thousand small objects stays a handful of API calls.

Reading with pandas, polars or duckdb

Bucket.uri hands the object store to the tools that already speak it, so data never has to pass through your process:

uri = rb.Bucket.from_name("forecasts").uri

pl.read_parquet(f"{uri}/2026/**/*.parquet")
duckdb.sql(f"SELECT * FROM '{uri}/2026/**/*.parquet'")

Inside a deployed run this works out of the box. From your own machine you will need credentials for the underlying bucket, so prefer get() or download() there.

Deleting

Deleting a bucket requires it to be empty, and the API refuses while anything remains. Empty it first:

b.delete_prefix("")   # or a narrower prefix
rebase bucket rm forecasts 2026/ --recursive
rebase bucket delete forecasts

Emptying happens from the client, one page at a time. Draining a large bucket inside a single API call would time out well before finishing, and the retry would resume against a half-emptied bucket.

A bucket still attached to a deployed function or app cannot be deleted either; the error names what is holding it.

CLI

rebase bucket create forecasts
rebase bucket list
rebase bucket ls forecasts 2026/ --delimiter /
rebase bucket put forecasts ./model.pkl models/model.pkl
rebase bucket download forecasts models/model.pkl ./model.pkl
rebase bucket uri forecasts
rebase bucket delete forecasts

See the rebase bucket reference for the full command list.

Bucket.uri and the bucket's console link both come from the API's bucket response, so rebase bucket uri forecasts prints the gs:// form and rebase tui can open the bucket in the Google Cloud console with o.

Permissions

Buckets use buckets:read and buckets:write. Both are granted to the Developer role, so the people deploying functions that use a bucket can also create and fill it. Viewers get buckets:read.

On this page