앞서 포스팅한 델타-뉴트럴 전략 스터디에 사용되었던 데이터를 어떻게 수집 했는지 남겨보려 합니다.
처음에 일봉 데이터를 이용해 백테스팅을 했더니, 선물 시장의 현실적인 변동성을 반영하지 못해서 백테스팅 결과가 너무 좋게 나오는 문제가 있었습니다. 그래서 시간봉을 이용해 백테스팅을 다시해 보았습니다.

– 현물 1시간봉, 선물 1시간봉, 펀딩 8시간만, 필요한 만큼만 받기
델타-뉴트럴(delta_neutral) 전략을 제대로 백테스트하려면 최소한 다음 3종류 데이터가 필요합니다.
- 현물(Spot) 1h OHLCV
- 선물(Perp Futures) 1h OHLCV
- 펀딩(Funding Rate) 8h: fundingTime, fundingRate
많은 예제에서 mark price(마크프라이스) 시계열까지 수집하지만, 실제로는 데이터 수집/관리 난이도만 올라가고, 백테스트목적(전략 비교/리밸런싱 구조 검증)에서는 선물 종가(close)를 mark price의 근사값으로 사용하는 것으로도 충분히 시작할 수 있습니다.
그래서 이번 스터디에는 mark price는 과감히 제외하고, 닥 3개의 CSV file만 생성해서 사용했습니다.
왜 mark price를 제외했나?
바이낸스 펀딩은 원칙적으로 mark price 기준으로 정산됩니다.
하지만 백테스트에서 mark price를 반드시 포함해야만 하는 것은 아닙니다.
펀딩 수익/비용의 “대략적인 크기”와 델타-뉴트럴 리밸런싱이 만드는 “계좌 변동성”을 보는 목적이라면, 선물 체결가(또는 종가) 기반으로 펀딩 명목을 근사해도 실무적으로 큰 문제가 없는 경우가 많습니다.
특히 초기 단계에서는 수집 실패율이 낮고, 코드가 단순하며, 실험 반복이 빠르다는 장점이 있습니다. 나중에 정밀도를 올릴 필요가 생기면 그때 markPriceKlines를 추가하는 편이 합리적입니다.
바이낸스 API의 “1000개 제한”은 어떻게 처리하나?
바이낸스의 kline API는 한 번에 최대 1000개 캔들만 돌려 줍니다.
그래서 6000시간(=6000개 캔들)을 받으려면 자동으로 여러 번 호출해서 이어 붙이는(pagination) 로직이 필요합니다.
이번 스크립트는 다음 방식으로 pagination을 구현했습니다.
- startTime부터 요청해서 최대 1000개를 받는다.
- 응답의 마지막 open_time을 기준으로 다음 요청의 startTime을 last_open_time+1로 갱신한다.
- endTime까지 반복한다.
덕분에 “스크립트를 한 번만 실행”해도 내부적으로 여러 번 API를 호출해 6000개를 채웠습니다.
사용법
기본 실행
python3 Binance_margin_data_gather_6000h_nomark.py호
환경변수 설정 변경
OUT_DIR=./data HOURS=6000 SYMBOL=BTCUSDT INTERVAL=1h SLEEP_SEC=0.2 \ python3 Binance_margin_data_gather_6000h_nomark.py
. HOURS : 몇 시간 분량을 받을지(시간봉 기준 캔들 개수)
. SLEEP_SEC : 요청 간 딜레이(레이트리밋 완화 목적)
. OUT_DIR : 저장 폴더
출력 파일
예를 들어 HOURS=6000, SIMBOL=BTCUSDT, INTERVAL=1h 이면
-
spot_BTCUSDT_1h_last6000h.csv -
futures_BTCUSDT_1h_last6000h.csv -
funding_BTCUSDT_last6000h.csv
이 3개면 델타-뉴트럴 전략의 1차 백테스트는 충분히 가능합니다.
백테스트에 어떻게 쓰나?
백테스트에서는 보통 다음과 같이 조합합니다.
- 시각 정렬 : spot 1h와 futures 1h는 open_time 기준으로 inner join
- 편딩: fundingTime(8시간 간격)을 시간축에 맵핑하여 해당 시점에만 펀딩을 반영
- mark price가 없으면 :
- funding_notional ≈ futures_close * short_qty로 근가
이 글에서는 데이터 수집기까지만 다뤘고,
기회가 된다면, 다음 글에서 ‘1시간 리밸런싱 델타-뉴트럴 백테스트’를 이어서 정리할 예정입니다.
제가 내린 결론은!
- 처음부터 모든 데이터를 다 모으기보다
전략을 검증하는 데 필요한 최소 데이터만 먼저 확보하는 것이 실험 속도를 크게 높인다. - mark price는 정밀도를 높이는 단계에서 추가해도 늦지 않다.
부족하지만, 제가 스터디에 이용한 도구에대한 포스팅을 마칠까합니다.
혹시 제안 사항이 있거나, 오류가 있으면 언제든 답글 부탁드립니다.
좋은 하루 되세요~~!! ^^
#!/usr/bin/env python3
# -*- coding: utf-8 -*-
"""
Binance 6000-hour (1h) Data Downloader (NO Mark Price)
------------------------------------------------------
This script downloads ONLY the datasets required to backtest an hourly delta-neutral strategy
without mark-price time series:
1) Spot 1h OHLCV (e.g., BTCUSDT spot)
- GET https://api.binance.com/api/v3/klines
2) Futures (USDⓈ-M Perpetual) 1h OHLCV (e.g., BTCUSDT perp)
- GET https://fapi.binance.com/fapi/v1/klines
3) Funding history (8h cadence; fundingTime, fundingRate)
- GET https://fapi.binance.com/fapi/v1/fundingRate
Why no mark price?
- Binance's funding is calculated on mark price, but many practical backtests approximate the funding
notional using futures trade price (close) when a mark-price series is not available.
- If you later need higher fidelity, you can add /fapi/v1/markPriceKlines, but this script intentionally
excludes it to keep data collection minimal and robust.
Outputs (default in current directory):
- spot_<SYMBOL>_<INTERVAL>_last<HOURS>h.csv
- futures_<SYMBOL>_<INTERVAL>_last<HOURS>h.csv
- funding_<SYMBOL>_last<HOURS>h.csv
Environment variables (optional):
- SYMBOL=BTCUSDT
- HOURS=6000
- INTERVAL=1h
- SLEEP_SEC=0.2
- OUT_DIR=./data
Run:
python3 Binance_margin_data_gather_6000h_nomark.py
"""
import os
import time
from dataclasses import dataclass
from datetime import datetime, timezone, timedelta
import requests
import pandas as pd
SPOT_BASE = "https://api.binance.com"
FUT_BASE = "https://fapi.binance.com"
@dataclass
class DownloadConfig:
symbol: str = "BTCUSDT"
interval: str = "1h"
hours: int = 6000
sleep_sec: float = 0.2
kline_limit: int = 1000
out_dir: str = "."
http_timeout: int = 10
def utc_now() -> datetime:
return datetime.now(timezone.utc)
def to_ms(dt: datetime) -> int:
return int(dt.timestamp() * 1000)
def ensure_dir(path: str):
os.makedirs(path, exist_ok=True)
def save_df(df: pd.DataFrame, path: str):
df.to_csv(path, index=False)
print(f"[OK] Saved: {path} rows={len(df)}")
def fetch_klines(base_url: str, endpoint: str, symbol: str, interval: str,
start_ms: int, end_ms: int,
limit: int = 1000, sleep_sec: float = 0.2, timeout: int = 10) -> pd.DataFrame:
"""Generic kline pagination (Binance limit=1000 per request)."""
url = base_url + endpoint
out = []
cur = start_ms
while cur < end_ms:
params = {
"symbol": symbol,
"interval": interval,
"startTime": cur,
"endTime": end_ms,
"limit": limit,
}
r = requests.get(url, params=params, timeout=timeout)
r.raise_for_status()
data = r.json()
if not data:
break
out.extend(data)
last_open_time = int(data[-1][0])
next_cur = last_open_time + 1
if next_cur <= cur:
break
cur = next_cur
time.sleep(sleep_sec)
if len(out) > 5_000_000:
raise RuntimeError("Too many rows; check interval or date range.")
cols = [
"open_time","open","high","low","close","volume",
"close_time","quote_asset_volume","num_trades",
"taker_buy_base","taker_buy_quote","ignore"
]
df = pd.DataFrame(out, columns=cols)
if df.empty:
return df
df["open_time"] = pd.to_datetime(df["open_time"], unit="ms", utc=True)
df["close_time"] = pd.to_datetime(df["close_time"], unit="ms", utc=True)
for c in ["open","high","low","close","volume","quote_asset_volume","taker_buy_base","taker_buy_quote"]:
df[c] = df[c].astype(float)
df["num_trades"] = df["num_trades"].astype(int)
return df
def fetch_funding_rate_history(symbol: str, start_ms: int, end_ms: int,
limit: int = 1000, sleep_sec: float = 0.2, timeout: int = 10) -> pd.DataFrame:
"""Futures funding history pagination by fundingTime (ms)."""
url = FUT_BASE + "/fapi/v1/fundingRate"
out = []
cur = start_ms
while cur < end_ms:
params = {
"symbol": symbol,
"startTime": cur,
"endTime": end_ms,
"limit": limit,
}
r = requests.get(url, params=params, timeout=timeout)
r.raise_for_status()
data = r.json()
if not data:
break
out.extend(data)
last_time = int(data[-1]["fundingTime"])
next_cur = last_time + 1
if next_cur <= cur:
break
cur = next_cur
time.sleep(sleep_sec)
if len(out) > 5_000_000:
raise RuntimeError("Too many funding rows; check date range.")
df = pd.DataFrame(out)
if df.empty:
return df
df["fundingTime"] = pd.to_datetime(df["fundingTime"], unit="ms", utc=True)
df["fundingRate"] = df["fundingRate"].astype(float)
df = df[["fundingTime", "fundingRate", "symbol"]].copy()
return df
def main():
cfg = DownloadConfig()
cfg.symbol = os.getenv("SYMBOL", cfg.symbol)
cfg.interval = os.getenv("INTERVAL", cfg.interval)
cfg.hours = int(os.getenv("HOURS", str(cfg.hours)))
cfg.sleep_sec = float(os.getenv("SLEEP_SEC", str(cfg.sleep_sec)))
cfg.out_dir = os.getenv("OUT_DIR", cfg.out_dir)
ensure_dir(cfg.out_dir)
end = utc_now()
start = end - timedelta(hours=cfg.hours)
start_ms = to_ms(start)
end_ms = to_ms(end)
print(f"[INFO] Download range UTC: {start.isoformat()} ~ {end.isoformat()} (hours={cfg.hours})")
print(f"[INFO] SYMBOL={cfg.symbol} INTERVAL={cfg.interval}")
spot = fetch_klines(
SPOT_BASE, "/api/v3/klines",
cfg.symbol, cfg.interval, start_ms, end_ms,
limit=cfg.kline_limit, sleep_sec=cfg.sleep_sec, timeout=cfg.http_timeout
)
save_df(spot, os.path.join(cfg.out_dir, f"spot_{cfg.symbol}_{cfg.interval}_last{cfg.hours}h.csv"))
fut = fetch_klines(
FUT_BASE, "/fapi/v1/klines",
cfg.symbol, cfg.interval, start_ms, end_ms,
limit=cfg.kline_limit, sleep_sec=cfg.sleep_sec, timeout=cfg.http_timeout
)
save_df(fut, os.path.join(cfg.out_dir, f"futures_{cfg.symbol}_{cfg.interval}_last{cfg.hours}h.csv"))
fund = fetch_funding_rate_history(
cfg.symbol, start_ms, end_ms,
limit=1000, sleep_sec=cfg.sleep_sec, timeout=cfg.http_timeout
)
save_df(fund, os.path.join(cfg.out_dir, f"funding_{cfg.symbol}_last{cfg.hours}h.csv"))
print("[DONE] Data collection finished (spot 1h / futures 1h / funding 8h).")
if __name__ == "__main__":
main()
이전 글>