ML meta-labeling & cloud export
Meta-labeling does not predict the market. It predicts your own signal: a rule-based primary strategy proposes every trade, and an ML filter learns which of those trades to actually take. Get the leakage controls right and it lifts expectancy; get them wrong and you ship a fantasy. This guide walks the full Lopez de Prado stack in edgekit — triple-barrier labels, purged walk-forward, a gradient-boosted filter, and a cloud-safe export a cTrader cBot runs with no scikit-learn.
import edgekit.ml itself needs only numpy + pandas, but the functions that fit models lazy-import their backend. Install what you use: pip install edgekit[ml] (scikit-learn + xgboost + lightgbm). A missing backend raises an actionable ImportError at make_model time, not at import.The pattern#
- Primary signal — a rule (e.g. the ORB breakout) proposes candidate trades.
- Labels — each candidate is labelled win/lose by a triple barrier that mirrors the engine's fills.
- Filter — a classifier learns
P(win)from causal features; you take a trade only whenP ≥ threshold. - Validation — purged, embargoed walk-forward so no label's outcome leaks into its own training set.
- Export — freeze the model to a tree blob a cBot evaluates in pure C# / Python, matching sklearn within 1e-6.
Step 1 — features#
ml.build_features builds causal feature families aligned to the bars — every feature at bar i uses only bars ≤ i. The higher-timeframe (mtf) family is shifted one completed bar so it never peeks.
import edgekit as ek
from edgekit.ml import build_features
bars = ek.data.resample_ohlcv(ek.data.load_bars("US100_M1.csv"), "4h")
X = build_features(bars, families=["price", "atr", "trend", "macd", "structure"])
# families default to all of price/atr/trend/macd/structure/volume/time/mtf/reversal/extra
# ('volume' is skipped automatically unless a tick_volume column is present)Step 2 — triple-barrier labels + uniqueness weights#
triple_barrierlabels each candidate bar in both directions. Its semantics mirror the backtest engine exactly, so labels and backtest agree: entry at the next bar's open, TP = entry + dir·rr·stop, SL = entry − dir·stop, a bar touching both counts as SL (pessimistic), and a vertical barrier at horizon_bars. y=1 iff TP is hit first.
from edgekit.ml import triple_barrier, uniqueness_weights, LabelConfig
# candidate bars: where your PRIMARY rule would fire. Here, a simple 20-bar high breakout.
import numpy as np
from edgekit import indicators as ind
from edgekit.core import lag
up, _ = ind.donchian(bars.high.to_numpy(), bars.low.to_numpy(), 20)
cand_mask = bars.high.to_numpy() >= lag(up, 1)
cfg = LabelConfig(rr=2.0, stop_atr_mult=1.0, atr_period=14, horizon_bars=48)
labels = triple_barrier(bars, cand_mask, cfg, cost=0.0)
# columns: entry_index, entry_time, exit_index, exit_time, dir, y, r, net, stop, outcome, weight
print(labels["y"].mean(), "base win rate over", len(labels), "candidates")Overlapping labels share future bars, so a naive fit over-counts redundant samples. uniqueness_weights gives each label its average-uniqueness weight (mean of 1/concurrency over its lifespan) — pass it as sample_weight when you fit.
w = uniqueness_weights(labels["entry_index"].to_numpy(),
labels["exit_index"].to_numpy(), n_bars=len(bars))Step 3 — purged walk-forward (the leakage firewall)#
This is the step that separates real meta-labeling from leakage theatre. A label's outcome is only known after its horizon plays out. If a training window contains a label whose future overlaps the test window, the model has seen the answer. Two mechanisms prevent it:
- Purge — drop any training label whose
exit_timeruns past the train-window boundary into the test period. - Embargo — additionally gap a buffer between train and test (≥ the label horizon) so nothing straddles the seam.
walk_forward_windows yields those windows as .iloc positions; use it as the out-of-sample backtest, training on the rolling past and testing on the next block.
from edgekit.ml import walk_forward_windows, WalkForwardConfig, make_model
from edgekit.ml import best_threshold, ev_at
wf = WalkForwardConfig(train_days=180, test_hours=24 * 30, embargo_hours=48) # embargo >= horizon
Xf = X.reindex(bars.index) # feature row per candidate entry
oos_r = []
for win in walk_forward_windows(labels, wf):
tr, te = labels.iloc[win["train_pos"]], labels.iloc[win["test_pos"]]
Xtr = Xf.iloc[tr["entry_index"].to_numpy()]
Xte = Xf.iloc[te["entry_index"].to_numpy()]
m = make_model("hgb", params={"max_iter": 40, "max_depth": 3})
m.fit(Xtr, tr["y"].to_numpy())
thr, _ = best_threshold(m.predict_proba(Xtr), tr["r"].to_numpy(), min_trades=10) # tune on TRAIN
mean_r, n = ev_at(m.predict_proba(Xte), te["r"].to_numpy(), thr) # score on TEST
oos_r.append((n, mean_r))Step 4 — the frozen meta-label#
Once validated, fit a final model and wrap it in a MetaLabeler. It takes a trade only when P(win) ≥ threshold (default 0.55). Built from a fitted estimator it exports the trees immediately and carries no sklearn dependency thereafter.
from edgekit.ml import MetaLabeler
m = make_model("hgb", params={"max_iter": 40, "max_depth": 3})
m.fit(Xf.iloc[labels["entry_index"].to_numpy()], labels["y"].to_numpy(), sample_weight=w)
meta = MetaLabeler(clf=m.estimator, threshold=0.55, feature_names=list(X.columns))
take = meta.take(X_live) # boolean take/skip per candidate — no sklearn needed
proba = meta.proba(X_live) # the P(win) scoresStep 5 — cloud-safe export for a cBot#
A cTrader cloud cBot cannot import scikit-learn. edgekit serialises a fitted HistGradientBoosting model to a plain tree-array blob and generates pure-Python or pure-C# inference from it. The reference forward pass matches sklearn's predict_proba to within 1e-6 — the round-trip guarantee.
from edgekit.ml import export_trees, tree_predict_proba, emit_csharp, emit_python
blob = export_trees(m.estimator) # "<baseline>#<tree>|<tree>..."
# verify the round-trip before you trust the export
import numpy as np
assert abs(tree_predict_proba(blob, Xf.to_numpy()) - m.predict_proba(Xf)).max() < 1e-6
cs = emit_csharp({"orb_meta": blob}, feature_names=list(X.columns)) # drop into a cAlgo cBot
py = emit_python({"orb_meta": blob}, feature_names=list(X.columns)) # or a pure-Python runtimeemit_csharp generates a self-contained static class (Predict(name, x), Threshold) embedding the blob and a pure-C# tree walk; emit_python generates the equivalent numpy-free Python module (predict(name, x) / take(name, x)). No model files, no runtime dependency — the whole classifier is a string in your source.
Step 6 — the leakage null test#
Before you believe any lift, prove the pipeline learns nothing from shuffled labels. shuffle_label_gate permutes the labels, refits, and confirms out-of-sample AUC collapses to ~0.5. If a model scores high on scrambled labels, there is a leak in the features or the split — fix it before reading any real result.
from edgekit.ml import shuffle_label_gate
gate = shuffle_label_gate(lambda: make_model("hgb", params={"max_iter": 40}),
Xf.iloc[labels["entry_index"].to_numpy()].to_numpy(),
labels["y"].to_numpy(), r=labels["r"].to_numpy(), n=10)
print(f"shuffled-label AUC {gate['mean_auc']:.3f} (want ~0.50) mean_ev {gate['mean_ev']:+.3f}R")shuffle_label_gate returns an AUC meaningfully above 0.5, your model is reading information it will not have live. Every real result must sit on top of a passing shuffle gate.Next#
- API · ml — config, labeling, splits, features, models, metrics, export.
- Causality — why purge + embargo exist.
- API · strategy — the ORB, the canonical meta-labeling primary signal.
- Proving an edge — the gauntlet the meta-labelled strategy still has to pass.