Create an autoregressive forecasting model and publish it to Huggingface

Build and backtest an autoregressive SE3 consumption forecaster, then deploy it to Rebase and publish the model version to Hugging Face.

This example starts with model development. It fetches daily SE3 consumption from the public eSett Open Data API, builds a simple autoregressive model, and compares it against a persistence baseline. Only after the local backtest looks reasonable do we deploy the model and publish the resulting Rebase model version to Hugging Face.

The target variable is SE3 total consumption in MWh from eSett's consumption endpoint. eSett exposes consumption by Metering Balance Area (MBA); SE3's MBA code is 10Y1001A1001A46L.

1. Install

Create a local Python environment:

uv venv .venv
source .venv/bin/activate
uv pip install "rebase-toolkit[huggingface]" numpy pandas requests scikit-learn

Configure Rebase now if you plan to deploy later:

rebase setup

Hugging Face authentication is not needed while developing the model. You only need it when you publish:

rebase connect huggingface

2. Fetch SE3 Consumption

Create se3_ar_forecast.py. Everything the model needs lives on one rebase.Predictor class, because rebase deploy ships only the class body and the module's imports. Module-level constants and helper functions are not deployed, so a predictor that calls them works locally and fails remotely with a NameError.

Start the file with the imports, the settings, and the data fetch:

from __future__ import annotations

import math
from typing import Any

import numpy as np
import pandas as pd
import requests
from sklearn.linear_model import Ridge
from sklearn.metrics import mean_absolute_error, mean_squared_error

import rebase


class SE3AutoregressiveConsumption(rebase.Predictor):
    project = "forecasting"
    name = "se3-autoregressive-consumption"
    description = "Daily SE3 consumption forecast using eSett Open Data and autoregressive lags."
    dependencies = [
        "numpy",
        "pandas",
        "requests",
        "scikit-learn",
    ]

    esett_api = "https://api.opendata.esett.com"
    se3_mba = "10Y1001A1001A46L"
    default_lags = [1, 2, 3, 7, 14, 28]

    def fetch_se3_consumption(self, start: str = "2025-01-01", end: str = "2026-05-31") -> pd.DataFrame:
        response = requests.get(
            f"{self.esett_api}/EXP15/Aggregate",
            params={
                "start": f"{start}T00:00:00.000Z",
                "end": f"{end}T00:00:00.000Z",
                "mba": self.se3_mba,
                "resolution": "Day",
            },
            headers={"Accept": "application/json"},
            timeout=30,
        )
        response.raise_for_status()

        rows = response.json()
        frame = pd.DataFrame(rows)
        if frame.empty:
            raise ValueError("eSett returned no SE3 consumption rows")

        frame["timestamp"] = pd.to_datetime(frame["timestampUTC"], utc=True)
        frame["consumption_mwh"] = -frame["total"].astype(float)
        return frame[["timestamp", "consumption_mwh"]].sort_values("timestamp").reset_index(drop=True)

The eSett API returns consumption values as negative MWh values. The example flips the sign so the model sees positive consumption.

3. Build Autoregressive Features

Add the feature-building methods to the class:

    def _feature_row(self, history: list[float], timestamp: pd.Timestamp, lags: list[int]) -> dict[str, float]:
        day = float(timestamp.dayofweek)
        row = {
            "dow_sin": math.sin(2.0 * math.pi * day / 7.0),
            "dow_cos": math.cos(2.0 * math.pi * day / 7.0),
            "trend": float(len(history)),
        }
        for lag in lags:
            row[f"lag_{lag}"] = float(history[-lag])
        return row

    def make_training_matrix(
        self,
        frame: pd.DataFrame,
        lags: list[int] | None = None,
    ) -> tuple[pd.DataFrame, pd.Series]:
        resolved_lags = lags or self.default_lags
        max_lag = max(resolved_lags)
        values = frame["consumption_mwh"].astype(float).tolist()
        timestamps = frame["timestamp"].tolist()

        rows: list[dict[str, float]] = []
        targets: list[float] = []
        for index in range(max_lag, len(values)):
            rows.append(self._feature_row(values[:index], timestamps[index], resolved_lags))
            targets.append(values[index])

        return pd.DataFrame(rows), pd.Series(targets, name="consumption_mwh")

    def fit_autoregressive_model(self, frame: pd.DataFrame, lags: list[int] | None = None) -> Ridge:
        features, target = self.make_training_matrix(frame, lags=lags)
        return Ridge(alpha=1.0).fit(features, target)

This is intentionally modest: lagged consumption explains most of the daily baseline, while day-of-week features capture a simple weekly shape.

4. Backtest Against Persistence

Before deploying anything, compare the autoregressive model to a persistence baseline. Persistence predicts that tomorrow will equal today.

Add the backtest method:

    def backtest(
        self,
        frame: pd.DataFrame,
        holdout_days: int = 60,
        lags: list[int] | None = None,
    ) -> dict[str, Any]:
        resolved_lags = lags or self.default_lags
        max_lag = max(resolved_lags)
        if len(frame) <= holdout_days + max_lag:
            raise ValueError("not enough data for the requested backtest")

        train = frame.iloc[:-holdout_days].reset_index(drop=True)
        test = frame.iloc[-holdout_days:].reset_index(drop=True)
        model = self.fit_autoregressive_model(train, lags=resolved_lags)

        history = train["consumption_mwh"].astype(float).tolist()
        ar_predictions: list[float] = []
        persistence_predictions: list[float] = []

        for _, row in test.iterrows():
            features = pd.DataFrame([self._feature_row(history, row["timestamp"], resolved_lags)])
            ar_prediction = float(model.predict(features)[0])
            persistence_prediction = float(history[-1])

            ar_predictions.append(ar_prediction)
            persistence_predictions.append(persistence_prediction)
            history.append(float(row["consumption_mwh"]))

        actual = test["consumption_mwh"].astype(float).to_numpy()
        ar_mae = mean_absolute_error(actual, ar_predictions)
        persistence_mae = mean_absolute_error(actual, persistence_predictions)

        return {
            "holdout_days": holdout_days,
            "ar_mae_mwh": float(ar_mae),
            "persistence_mae_mwh": float(persistence_mae),
            "ar_rmse_mwh": float(mean_squared_error(actual, ar_predictions) ** 0.5),
            "persistence_rmse_mwh": float(mean_squared_error(actual, persistence_predictions) ** 0.5),
            "mae_improvement_vs_persistence_pct": float(100.0 * (persistence_mae - ar_mae) / persistence_mae),
        }

Add a module-level instance at the end of the file, outside the class, so scripts can import it:

model = SE3AutoregressiveConsumption()

Run a local backtest:

python - <<'PY'
import json

from se3_ar_forecast import model

data = model.fetch_se3_consumption()
metrics = model.backtest(data)
print(json.dumps(metrics, indent=2))
PY

At this point you are still just developing the model. If persistence wins, change the lags, add calendar features, use a longer history window, or test a different model class before deploying.

5. Add the Prediction Entry Point

Once the local backtest is acceptable, add the forecast and the predict method that Rebase calls. Place them inside the class, above the model = ... line:

    def forecast_next_days(
        self,
        frame: pd.DataFrame,
        horizon_days: int = 7,
        lags: list[int] | None = None,
    ) -> list[dict[str, Any]]:
        resolved_lags = lags or self.default_lags
        model = self.fit_autoregressive_model(frame, lags=resolved_lags)

        history = frame["consumption_mwh"].astype(float).tolist()
        timestamp = frame["timestamp"].iloc[-1]
        forecast: list[dict[str, Any]] = []

        for _ in range(horizon_days):
            timestamp = timestamp + pd.Timedelta(days=1)
            features = pd.DataFrame([self._feature_row(history, timestamp, resolved_lags)])
            value = float(model.predict(features)[0])
            history.append(value)
            forecast.append(
                {
                    "timestamp": timestamp.isoformat(),
                    "consumption_mwh": value,
                }
            )

        return forecast

    def predict(
        self,
        start: str = "2025-01-01",
        end: str = "2026-05-31",
        horizon_days: int = 7,
        holdout_days: int = 60,
    ) -> dict[str, Any]:
        frame = self.fetch_se3_consumption(start=start, end=end)
        return {
            "target": "SE3 daily total consumption",
            "unit": "MWh",
            "training_start": frame["timestamp"].iloc[0].isoformat(),
            "training_end": frame["timestamp"].iloc[-1].isoformat(),
            "backtest": self.backtest(frame, holdout_days=holdout_days),
            "forecast": self.forecast_next_days(frame, horizon_days=horizon_days),
        }

predict uses literal defaults. A default that names a module-level constant would be evaluated when the deployed class is defined, where that constant does not exist.

Run it locally:

python - <<'PY'
from se3_ar_forecast import model

result = model.predict(horizon_days=3)
print(result["backtest"])
print(result["forecast"])
PY

6. Deploy to Rebase

Deploy the model to dev without involving Hugging Face yet:

rebase deploy se3_ar_forecast.py --name model --env dev

rebase deploy creates or updates the Rebase model and deploys the current version. Without --env it targets your active environment, which is dev unless you switched. Inspect the model version and environment pointer:

rebase model versions se3-autoregressive-consumption --project forecasting
rebase model deployments se3-autoregressive-consumption --project forecasting

Validate the deployed model:

rebase model run se3-autoregressive-consumption \
  --project forecasting \
  --env dev \
  --param horizon_days=3

7. Publish to Hugging Face

Publishing is a separate step after the model and backtest flow are working. First, create a Hub README as README.hf.md:

---
library_name: scikit-learn
tags:
  - forecasting
  - time-series
  - electricity-consumption
  - se3
  - esett
---

# SE3 Autoregressive Consumption Forecast

Daily SE3 total consumption forecast trained on eSett Open Data.

The model uses lagged daily consumption, day-of-week features, and a linear Ridge regressor. The accompanying `backtest.json` compares the autoregressive forecast against a persistence baseline.

The Rebase SDK also publishes `rebase_model.json` with model-version provenance: the Rebase model version, fingerprint, source hash and Git metadata. The Hub commit itself is recorded on the Rebase publication.

Then create a small artifact folder with the data window, backtest summary, and Hub README:

import json
from pathlib import Path

from se3_ar_forecast import model

import rebase


HUGGINGFACE_README = Path("README.hf.md")


def write_huggingface_artifacts() -> Path:
    artifact_dir = Path("artifacts/se3-autoregressive-consumption")
    artifact_dir.mkdir(parents=True, exist_ok=True)

    frame = model.fetch_se3_consumption()
    frame.to_csv(artifact_dir / "training_data.csv", index=False)

    metrics = model.backtest(frame)
    (artifact_dir / "backtest.json").write_text(json.dumps(metrics, indent=2) + "\n", encoding="utf-8")
    (artifact_dir / "README.md").write_text(HUGGINGFACE_README.read_text(encoding="utf-8"), encoding="utf-8")
    return artifact_dir


deployed = model.deploy(
    environment="dev",
    huggingface=rebase.HuggingFacePublishConfig(
        repo_id="your-hf-org/se3-autoregressive-consumption",
        repo_type="model",
        private=False,
        artifact_path=write_huggingface_artifacts(),
        sync_source_git=True,
    ),
)

print("huggingface_publication:", deployed.data["huggingface_publication"])

Save that as publish_huggingface.py and run:

python publish_huggingface.py

The Hub repo receives the artifact folder plus Rebase provenance. Rebase records the Hub repo, revision, commit SHA, visibility, and the model version that produced the publication.

Outside a clean Git checkout the SDK skips the Git link automatically, and the publication is recorded as published instead of in_sync. Set sync_source_git=False to always skip it:

rebase.HuggingFacePublishConfig(
    repo_id="your-hf-org/se3-autoregressive-consumption",
    artifact_path=write_huggingface_artifacts(),
    sync_source_git=False,
)

8. Promote to Production

Promote the same immutable model version to staging:

rebase model promote se3-autoregressive-consumption \
  --project forecasting \
  --from dev \
  --to staging

Request production approval:

rebase model request-promotion se3-autoregressive-consumption \
  --project forecasting \
  --from staging \
  --to prod \
  --reason "Backtest passed against persistence and Hugging Face publication is recorded"

Approve the request:

rebase model approve-promotion <promotion-request-id> \
  --reason "Approved for production"

Promote to production:

rebase model promote se3-autoregressive-consumption \
  --project forecasting \
  --from staging \
  --to prod \
  --promotion-request-id <promotion-request-id>

Run production:

rebase model run se3-autoregressive-consumption \
  --project forecasting \
  --env prod \
  --param horizon_days=7

9. Inspect Lineage

List model versions and environment pointers:

rebase model versions se3-autoregressive-consumption --project forecasting
rebase model deployments se3-autoregressive-consumption --project forecasting

The Rebase model publication links the immutable model version to the Hugging Face repo and Hub commit. The Hub repo contains README.md, training_data.csv, backtest.json, rebase_model.json, and rebase_source.py.

On this page