← TokenPylon

API documentation

Public, free, no auth required. Rate limit: 60 req/min per IP.

Cite the original sources when reusing the data (each row carries its source_url). Epoch AI data is CC-BY 4.0; SEC EDGAR is public domain; press extractions cite back to the original article.

Prices

Hosted-API prices per (provider, model_key) from LiteLLM + scraped provider sites. Updated daily.

GET/api/v1/prices/current

All current prices across 2,278 models.

Returns input/output price for every (provider, model_key) with its unit: per_million_tokens for LLMs, per_image for image generation/editing, per_second for transcription (input) and speech/video output, per_million_characters for text-to-speech. Filters: ?provider=, ?model_key=, ?mode=chat|image_generation|audio_speech|audio_transcription|video_generation, ?unit=.

Example
curl 'https://tokenpylon.vercel.app/api/v1/prices/current?provider=anthropic'
Response shape
{"prices":[{"model_key":"claude-sonnet-4-5","provider":"anthropic","input_price_per_million":3,"output_price_per_million":15}]}
GET/api/v1/prices/on-date

Price snapshot for a specific date.

Time-travel query. Returns prices as they were on the given date.

Example
curl 'https://tokenpylon.vercel.app/api/v1/prices/on-date?date=2026-01-01&model_key=claude-sonnet-4-5'
Response shape
{"prices":[...]}
GET/api/v1/trends

Daily price index: median, quartiles and cheapest paid chat price per day, by scope.

One point per indexed day for ?scope=all (default), provider:<name>, family:<model_family> or decide:<task>. Each point carries the listing count, sources_covered/sources_total, interpolated p25/median/p75 input and output prices in USD per million tokens, the cheapest price, and the listings cut or raised that day. A day a source did not poll is absent for its listings, never carried forward. decide:<task> points hold that day's balanced /decide recommendation in detail. ?since= and ?until= (YYYY-MM-DD, default last 90 days, at most 400); ?list=1 returns the scopes indexed this week. Computed nightly from the price ledger since 2026-05-22.

Example
curl 'https://tokenpylon.vercel.app/api/v1/trends?scope=provider:openai&since=2026-08-01'
Response shape
{"scope":"provider:openai","since":"2026-08-01","until":"2026-09-22","unit":"usd_per_million_tokens","points":[{"day":"2026-08-01","listings":41,"sources_covered":3,"sources_total":3,"median_input":1.25,"median_output":10,"p25_input":0.4,"p75_input":2.5,"p25_output":1.6,"p75_output":10,"min_input":0.05,"min_output":0.4,"cuts":0,"raises":0,"detail":null}]}

Cost optimization

Pick the cheapest / greenest way to run a workload across hosted-API, hosted-GPU, and self-deploy.

POST/api/v1/cost/best-deploy

Cheapest end-to-end way to run a workload.

Returns up to ~15 deployment options ranked by cost. Set useLiveCarbon=true to overlay current grid carbon onto self-deploy options.

Example
curl -X POST 'https://tokenpylon.vercel.app/api/v1/cost/best-deploy' -H 'content-type: application/json' \
  -d '{"family":"llama-3-1-70b-instruct","inputTokens":1000000,"outputTokens":200000,"useLiveCarbon":true}'
Response shape
{"cheapest_by_cost":{"provider":"fireworks_ai","cost_per_million_tokens":0.100,...},"greenest_by_emissions":{...},"options":[...]}
POST/api/v1/schedule/best-window

Time-shift scheduler: when + where to run for lowest carbon.

Given a workload + deadline window in hours, recommends the (city, hour) tuple that minimizes emissions using historical hour-of-day carbon patterns.

Example
curl -X POST 'https://tokenpylon.vercel.app/api/v1/schedule/best-window' -H 'content-type: application/json' \
  -d '{"family":"llama-3-1-70b-instruct","inputTokens":1000000,"outputTokens":200000,"deadlineHours":24}'
Response shape
{"recommended":{"city_code":"us-quincy","suggested_hour_utc":6,"estimated_emissions_kg":0.11,...},"baseline_if_run_now":{...},"reduction_pct":89}

Infrastructure

Where the compute lives — provider locations, datacenter announcements, electricity.

GET/api/v1/provider-locations

GPU provider datacenter locations.

46 locations across Vultr (33) + Modal (13). Joined to electricity_rates.city_code where the city has carbon data.

Example
curl 'https://tokenpylon.vercel.app/api/v1/provider-locations?has_carbon=true'
Response shape
{"locations":[{"provider":"vultr","city_name":"Frankfurt","city_code":"eu-frankfurt"}]}
GET/api/v1/electricity

Per-city electricity rates + carbon intensity + PUE.

19 cities covered. Each row has retail rate, datacenter PUE estimate, and annual carbon intensity from eGRID / Ember.

Example
curl 'https://tokenpylon.vercel.app/api/v1/electricity'
Response shape
{"cities":[{"city_code":"us-quincy","rate_usd_kwh":0.05,"carbon_intensity_gco2_kwh":286.5}]}

Market intelligence

The AI infrastructure economy — chips, datacenters, training runs, capex, sanctions.

GET/api/v1/market

Single-shot dashboard summary across all 13 datasets.

Counts + top-5 for each table. Use this to discover the other endpoints.

Example
curl 'https://tokenpylon.vercel.app/api/v1/market'
Response shape
{"capex":{...},"gpu_clusters":{"count":680,...},"training_runs":{...},"bis_entity_list":{"count":3399,...}}
GET/api/v1/market/capex

Hyperscaler quarterly capex from SEC EDGAR.

5 hyperscalers (MSFT/GOOGL/META/AMZN/AVGO). ~10 years of history. Filters: ?ticker=, ?form=10-Q|10-K, ?since=.

Example
curl 'https://tokenpylon.vercel.app/api/v1/market/capex?ticker=MSFT&since=2024-01-01'
Response shape
{"rows":[{"ticker":"MSFT","period_end":"2026-03-31","capex_usd":30900000000}]}
GET/api/v1/market/tsmc

TSMC monthly net revenue.

Source: SEC EDGAR 6-K filings (TSMC's own site is Cloudflare-blocked). Includes MoM%, YoY%, YTD.

Example
curl 'https://tokenpylon.vercel.app/api/v1/market/tsmc'
Response shape
{"rows":[{"month":"2026-04-01","revenue_twd_m":410725,"yoy_pct":17.5}]}
GET/api/v1/market/clusters

Frontier GPU clusters with lat/lon, MW, owner.

From Epoch AI. ?current=true filters to non-superseded snapshots (avoids triple-counting xAI Colossus Phase 1+2+3).

Example
curl 'https://tokenpylon.vercel.app/api/v1/market/clusters?current=true&sort=h100&limit=10'
Response shape
{"rows":[{"cluster_name":"xAI Colossus Memphis Phase 3","owner":"xAI","h100_equivalents":275796}]}
GET/api/v1/market/training-runs

Notable ML training runs with compute + cost.

From Epoch AI's Notable Models dataset. 1,074 rows back to 1950. Filters: ?org=, ?frontier=true, ?since=.

Example
curl 'https://tokenpylon.vercel.app/api/v1/market/training-runs?frontier=true&sort=cost&limit=5'
Response shape
{"rows":[{"model_name":"Grok 4","organization":"xAI","training_cost_2023_usd":388000000}]}
GET/api/v1/market/hardware

AI accelerator specs (FP4..FP64, memory, TDP).

175 chips: NVIDIA, AMD, Google TPU, Amazon Trainium, Meta MTIA, Cerebras, Groq, Graphcore, Huawei Ascend.

Example
curl 'https://tokenpylon.vercel.app/api/v1/market/hardware?manufacturer=NVIDIA&sort=date'
Response shape
{"rows":[{"hardware_name":"NVIDIA GB300","release_date":"2025-08-22","fp8_flops":5e15,"tdp_watts":1400}]}
GET/api/v1/market/companies

AI lab financials (revenue, valuation, staff).

9 AI labs from Epoch AI. Sort by ?sort=valuation|revenue|staff|funding.

Example
curl 'https://tokenpylon.vercel.app/api/v1/market/companies?sort=revenue'
Response shape
{"rows":[{"company_name":"Anthropic","valuation_usd_latest":380000000000,"annual_revenue_usd_latest":30000000000}]}
GET/api/v1/market/datacenters

Real-time datacenter buildout announcement firehose.

DCK + DCD + Bisnow RSS feeds, LLM-extracted (operator, city, MW, online_date). Default returns only extracted buildouts; use ?status=all.

Example
curl 'https://tokenpylon.vercel.app/api/v1/market/datacenters?operator=Microsoft'
Response shape
{"rows":[{"title":"...","operator":"...","city_name":"...","power_mw":120}]}
GET/api/v1/market/ppas

Hyperscaler renewable PPA + nuclear restart firehose.

MSFT / Google / Meta / Amazon press feeds, LLM-classified for energy deals.

Example
curl 'https://tokenpylon.vercel.app/api/v1/market/ppas?buyer=Microsoft'
Response shape
{"rows":[{"buyer":"Microsoft","generation_type":"nuclear-restart","capacity_mw":835}]}
GET/api/v1/market/top500

Top500 supercomputer rankings biannually since 1993.

Last 10 years by default. Filters: ?list=2025-11-01, ?country=, ?gpu=H100.

Example
curl 'https://tokenpylon.vercel.app/api/v1/market/top500'
Response shape
{"rows":[{"rank":1,"system_name":"El Capitan","rmax_pflops":1809,"gpu_chip":"MI300A"}]}
GET/api/v1/market/export-controls

US Consolidated Screening List (Entity List + 8 other lists).

25k+ entities. Filters: ?list=entity|meu|cmic|sdn|all, ?country=CN, ?since=2024-01-01.

Example
curl 'https://tokenpylon.vercel.app/api/v1/market/export-controls?list=entity&country=CN&since=2024-01-01'
Response shape
{"rows":[{"name":"...","start_date":"2025-10-08","addresses":"Beijing..."}]}
GET/api/v1/market/interconnection

US grid interconnection queues from PJM/MISO/NYISO/CAISO/ERCOT.

19,383 projects. Filters: ?iso=, ?state=, ?fuel=, ?status=, ?min_mw=.

Example
curl 'https://tokenpylon.vercel.app/api/v1/market/interconnection?iso=ERCOT&min_mw=1000&fuel=GAS'
Response shape
{"rows":[{"project_id":"30INR0071","capacity_mw":1660,"county":"Franklin","fuel_type":"GAS"}]}

Savings workflow

Audit what you spend, find the same model cheaper, consider other models with an honest requires_eval flag, deploy a routing policy, then verify the saving on cost per successful task. Every number is a list price with an observation time; nothing here knows your negotiated rate.

POST/api/v1/cost/quote

Itemised workload quote: uncached input, cache reads/writes, output, batch rates, long-context tiers.

Body: { model, provider?, inputTokens, outputTokens, cachedTokens?, cacheWriteTokens?, batch?, calls? }. inputTokens excludes cached tokens. Long-context tiers apply automatically when the prompt exceeds a published threshold. Returns lines with the ledger row each rate came from, unknowns (every assumption), complete=false when a rate was missing, price_observed_at and a 24h quote_expires_at.

Example
curl -X POST 'https://tokenpylon.vercel.app/api/v1/cost/quote' -H 'content-type: application/json' -d '{"model":"claude-sonnet-4-5","inputTokens":50000,"cachedTokens":200000,"outputTokens":4000,"batch":true,"calls":1000}'
Response shape
{"model_key":"claude-sonnet-4-5","provider":"anthropic","per_call":0.1,"total":100,"complete":true,"applied":{"batch":true,"context_tier":"above-200k","caching":true},"unknowns":[],"lines":[{"kind":"input","tokens":50000,"rate_per_million":3,"cost":0.15,"from":"claude-sonnet-4-5:batch"}],"provenance":{"price_observed_at":"...","quote_expires_at":"..."}}
POST/api/v1/cost/audit

Spend audit: your usage repriced at today's rates, plus the cheapest other host selling the same model.

Body: { lines: [{ model, provider?, inputTokens, outputTokens, cachedTokens?, calls?, actualCostUsd?, label? }], period?, allowedProviders? }. Same-model-other-host is the one saving that needs no quality judgement. Lines that can't be fully priced come back under unresolved, never zeroed. Batch options are shown where a batch rate is published.

Example
curl -X POST 'https://tokenpylon.vercel.app/api/v1/cost/audit' -H 'content-type: application/json' -d '{"period":"2026-08","lines":[{"model":"meta-llama/llama-3.3-70b-instruct","provider":"together_ai","inputTokens":8000,"outputTokens":1200,"calls":50000}]}'
Response shape
{"totals":{"baseline":492,"cheapest_same_model":60,"savings":432,"savings_pct":87.8},"lines":[{"model_key":"...","provider":"together_ai","cheapest_same_model":{"provider":"deepinfra","total":60,"savings_pct":87.8}}],"unresolved":[]}
POST/api/v1/decide/alternatives

Cheaper models for the same task, within a quality tolerance of the model you use. Every row requires_eval.

Body: { model, task?, qualityTolerance?, allowedProviders?, minTokensPerSec?, maxTtftMs?, openWeightsOnly?, requires?, limit? }. The bar is YOUR model's benchmark score on the task metric, not a fixed floor: LiveBench coding for coding, LiveBench average for chat and cheap-batch, SimpleBench for reasoning. A model LiveBench has not scored is compared on Aider polyglot (frozen since October 2025) against other Aider results only; baseline.metric names the one used. Capability requirements (tools, json, vision, contextTokens) are echoed back as not_enforced: the catalog carries no capability data yet.

Example
curl -X POST 'https://tokenpylon.vercel.app/api/v1/decide/alternatives' -H 'content-type: application/json' -d '{"model":"gpt-4.1","task":"coding","qualityTolerance":0.1}'
Response shape
{"baseline":{"model_key":"gpt-4.1","score":0.52,"cost_per_run_usd":9.9},"alternatives":[{"model_key":"deepseek/deepseek-chat","score":0.50,"cost_per_run_usd":0.88,"cost_vs_baseline_pct":-91,"requires_eval":true,"evidence":"aider-polyglot 50.0% vs your 52.0%; ..."}]}
POST/api/v1/cost/verify

Did the switch save money? Cost per successful task, baseline vs candidate, with failures, retries and fallbacks counted.

Body: { baseline: [lines], candidate: [lines] } where a line is usage plus calls, successes/failures, retries, fallbacks and optionally actualCostUsd. A cheaper model that fails more is not a saving. Stateless: nothing is stored.

Example
curl -X POST 'https://tokenpylon.vercel.app/api/v1/cost/verify' -H 'content-type: application/json' -d '{"baseline":[{"model":"gpt-4.1","inputTokens":8000,"outputTokens":1000,"calls":1000,"failures":20}],"candidate":[{"model":"deepseek/deepseek-chat","inputTokens":8000,"outputTokens":1100,"calls":1000,"failures":45,"retries":30}]}'
Response shape
{"baseline":{"cost_per_successful_task":0.0244},"candidate":{"cost_per_successful_task":0.0031},"delta":{"pct":-87.3,"verdict":"saving"},"notes":["candidate success rate is 2.5 points lower; ..."]}

Agents & routing

The routing brain as a remote MCP server, plus a LiteLLM router config priced from live data. Your gateway does the routing; TokenPylon keeps its price table honest.

POST/api/mcp

Remote MCP server (Streamable HTTP). Tools: audit_spend, quote_job, cheaper_alternatives, verify_savings, decide, best_provider, estimate_cost, compare_cost, project_spend, families, price_changes, media_prices, litellm_router_config, watch_model.

Attach it once and your agent can pick a model for a job, find the cheapest provider under speed/latency constraints, price the work, and export a LiteLLM config. Stateless, no auth, same rate limit as the REST API. Claude Code: claude mcp add --transport http tokenpylon https://tokenpylon.vercel.app/api/mcp

Example
curl -X POST 'https://tokenpylon.vercel.app/api/mcp' -H 'content-type: application/json' -H 'accept: application/json, text/event-stream' -d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}'
Response shape
{"jsonrpc":"2.0","id":1,"result":{"tools":[{"name":"decide",...},{"name":"best_provider",...}]}}
GET/api/v1/route/litellm

LiteLLM Router model_list with cost-based routing, priced from live data.

?family=a,b names families; ?task=coding&priority=balanced&top=5 lets /decide pick them; ?providers= restricts deployments; ?format=yaml returns a drop-in config.yaml. Deployments carry input_cost_per_token / output_cost_per_token so LiteLLM's cost-based-routing and spend tracking see real prices. Add your own api_key / api_base per provider. Regenerate daily.

Example
curl 'https://tokenpylon.vercel.app/api/v1/route/litellm?family=llama-3-3-70b-instruct&format=yaml'
Response shape
model_list:
  - model_name: llama-3-3-70b-instruct
    litellm_params:
      model: deepinfra/meta-llama/Llama-3.3-70B-Instruct-Turbo
      input_cost_per_token: 1.0e-7
      output_cost_per_token: 3.2e-7
router_settings:
  routing_strategy: cost-based-routing

Collector

A small signed binary you install once. Your apps point their base URL at it (or you run them through it); every call is forwarded unchanged and its usage metadata recorded per attempt — model asked for and served, billing provider and gateway, tokens by kind, provider-reported cost, timing, outcome — never messages, never keys. A local page shows your usage priced at today's list rates and the one switch worth making; the catalog gets prices verified by real workloads.

CLItokenpylon up

Installs the collector as a login service on 127.0.0.1:4141 and opens the local usage page: totals, tokens and cost per token by source (proxy, LiteLLM, Claude Code, Codex), spend at list, and what other contributors pay on the same models. `tokenpylon mcp` is an MCP server any agent can query (live sessions and context, quota meters, usage, priced breakdown); `tokenpylon hook context` warns a harness past 100k tokens of context. Any client can tag its calls with X-Tokenpylon-Session, X-Tokenpylon-Agent, X-Tokenpylon-Project and X-Tokenpylon-Harness headers (stripped before forwarding) to group on the page by session and agent; without them the User-Agent names the harness and calls less than 15 minutes apart form a session. Claude Code and Codex usage is read from their own history files automatically; for everything else, `tokenpylon connect shell` (or claude-code, codex) or `tokenpylon run -- <your command>`.

Base-URL adapters: /openai/v1, /anthropic, /gemini, /openrouter/v1, /deepseek/v1, /mistral, /groq/v1, ... and /proxy/<api-host> for anything OpenAI-shaped. Streaming passes through unbuffered; usage is read from the provider's final event; a client that disconnects tears the upstream call down and the attempt is recorded as cancelled. `tokenpylon import litellm <spend-logs.json>` and `tokenpylon import openrouter <activity.csv>` bring in history; the LiteLLM hook (pip install tokenpylon-litellm) records every attempt the proxy makes. Sharing call metadata with tokenpylon.com is on by default and disclosed at first run: `tokenpylon sharing off` or TOKENPYLON_SHARE=0 turns it off, and the local page works either way. `tokenpylon sharing show` prints the exact fields. Install: npm i -g tokenpylon, brew install tokenpylon/tap/tokenpylon, or the container image for a sidecar.

Example
npm i -g tokenpylon && tokenpylon up && tokenpylon connect shell
Response shape
TokenPylon sends call metadata to tokenpylon.com by default; never messages or API keys.
We use it for your audit and shared price evidence, and retain it indefinitely.
Turn sharing off: tokenpylon sharing off — or set TOKENPYLON_SHARE=0.
collector installed (~/Library/LaunchAgents/com.tokenpylon.collector.plist)
dashboard: http://127.0.0.1:4141/
POST/api/v1/telemetry/audit

What the local page calls: your usage totals per (provider, model, host) priced at list, the cheapest same-model host for that exact token mix, the one opportunity worth acting on, and the catalog keys to watch.

Body: { period?, lines: [{ provider, model, served_host, source?, calls, input_tokens, cached_tokens, cache_write_tokens, cache_write_1h_tokens?, output_tokens, reported_cost_usd?, reported_calls? }] }. cache_write_1h_tokens is the part of the writes at the one-hour cache TTL, priced at its own tier. Stateless and unauthenticated: nothing is stored, so it works with sharing off. Only a provider-reported charge counts as a bill; LiteLLM's cost-map figure is calculated, not billed, and never verifies a price. An opportunity needs at least $1 and 5% to be named.

Example
curl -X POST 'https://tokenpylon.vercel.app/api/v1/telemetry/audit' -H 'content-type: application/json' -d '{"lines":[{"provider":"together_ai","model":"meta-llama/Llama-3.3-70B-Instruct-Turbo","served_host":"together_ai","calls":5000,"input_tokens":5000000,"cached_tokens":0,"cache_write_tokens":0,"output_tokens":1000000}]}'
Response shape
{"totals":{"baseline":5.28,"cheapest_same_model":0.82,"savings":4.46,"savings_pct":84.5,"reported":0},"opportunity":{"kind":"switch_host","headline":"$4.46 (84%) on meta-llama/Llama-3.3-70B-Instruct-Turbo by calling it at deepinfra",...},"watch":["meta-llama/Llama-3.3-70B-Instruct-Turbo"]}
GET/api/v1/telemetry/aggregates

The return feed: what contributed real traffic says about a model's hosts, published only above 5 installs, 1,000 calls, and no install over half.

?model=&provider=&days=30. Per (provider, model, host, cache band, context band): what providers actually charged per million tokens where they report cost, the catalog's quote for the same usage, the gap between them, error rate, and observed time-to-first-token and latency. Raw events are POSTed by the collector to /api/v1/telemetry/usage using a TokenPylon account API key (Authorization: Bearer). Raw records are isolated by account; public cohorts remain shared. Configure the recorder with tokenpylon auth set --stdin. Uploads stay queued until a key is configured. Events use a strict field allowlist (schema v2: per attempt, with logical call id, requested vs returned model, gateway vs billing provider, and an amount basis on every dollar figure) and kept indefinitely; cohorts are rolled up nightly.

Example
curl 'https://tokenpylon.vercel.app/api/v1/telemetry/aggregates?model=deepseek-v4-pro'
Response shape
{"count":2,"aggregates":[{"provider":"deepseek","model":"deepseek-v4-pro","served_host":"deepseek","calls":4120,"installs":9,"reported_usd_per_million":1.48,"quoted_usd_per_million":1.51,"reported_vs_quoted_pct":-2.0,"error_rate":0.004,"p50_ttft_ms":610}]}

Verified prices

A scraped price is a claim. Every night TokenPylon makes one tiny real call against a rolling slice of listings with its own keys, reads the bill back where the provider allows it, and records billed vs quoted. Quotes carry the result.

GET/api/v1/health/verification

Coverage of the verification job: listings verified by method, recent mismatches, spend to date, the caps in force.

Methods: response_cost (the provider returned dollars, e.g. OpenRouter), balance_delta (a prepaid balance moved by exactly the call), usage_only (token accounting checked, bill not observable in-band), unreachable. A quote's `verification` field carries the latest row for that listing, or null when never tested. Spend is capped per run and per month and every call is one ledger row.

Example
curl 'https://tokenpylon.vercel.app/api/v1/health/verification'
Response shape
{"enabled":true,"coverage":{"listings_verified":412,"reachable":398,"by_method":{"response_cost":260,"balance_delta":90,"usage_only":48,"unreachable":14}},"spend":{"month_to_date_usd":0.41},"mismatches":[{"model_key":"...","provider":"...","delta_pct":6.2}]}

Health & status

Internal health endpoints exposed publicly so anyone can audit data freshness.

GET/api/v1/health/freshness

Days-since-last-success per data source.

Surfaces stale scrapers before they affect any chart.

Example
curl 'https://tokenpylon.vercel.app/api/v1/health/freshness'
Response shape
{"summary":{"fresh":7,"stale":0,"dead":0},"sources":[{"source":"litellm","status":"fresh","days_since_last_success":0.9}]}
GET/api/v1/health/coverage

Per-source row counts and last-update times.

How many rows each scraper has loaded and when it last successfully ran.

Example
curl 'https://tokenpylon.vercel.app/api/v1/health/coverage'
Response shape
{"sources":[{"source":"epoch-gpu-clusters","row_count":680,"last_run_at":"..."}]}
GET/api/v1/health/cross-check

Disagreement between independent sources.

Where two sources report the same data point, surface which pairs disagree most. Used to find suspect numbers.

Example
curl 'https://tokenpylon.vercel.app/api/v1/health/cross-check'
Response shape
{"rows":[]}