Public, free, no auth required. Rate limit: 60 req/min per IP.
Cite the original sources when reusing the data (each row carries its source_url). Epoch AI data is CC-BY 4.0; SEC EDGAR is public domain; press extractions cite back to the original article.
Hosted-API prices per (provider, model_key) from LiteLLM + scraped provider sites. Updated daily.
/api/v1/prices/currentAll current prices across 2,278 models.
Returns input/output price for every (provider, model_key) with its unit: per_million_tokens for LLMs, per_image for image generation/editing, per_second for transcription (input) and speech/video output, per_million_characters for text-to-speech. Filters: ?provider=, ?model_key=, ?mode=chat|image_generation|audio_speech|audio_transcription|video_generation, ?unit=.
curl 'https://tokenpylon.vercel.app/api/v1/prices/current?provider=anthropic'
{"prices":[{"model_key":"claude-sonnet-4-5","provider":"anthropic","input_price_per_million":3,"output_price_per_million":15}]}/api/v1/prices/on-datePrice snapshot for a specific date.
Time-travel query. Returns prices as they were on the given date.
curl 'https://tokenpylon.vercel.app/api/v1/prices/on-date?date=2026-01-01&model_key=claude-sonnet-4-5'
{"prices":[...]}/api/v1/trendsDaily price index: median, quartiles and cheapest paid chat price per day, by scope.
One point per indexed day for ?scope=all (default), provider:<name>, family:<model_family> or decide:<task>. Each point carries the listing count, sources_covered/sources_total, interpolated p25/median/p75 input and output prices in USD per million tokens, the cheapest price, and the listings cut or raised that day. A day a source did not poll is absent for its listings, never carried forward. decide:<task> points hold that day's balanced /decide recommendation in detail. ?since= and ?until= (YYYY-MM-DD, default last 90 days, at most 400); ?list=1 returns the scopes indexed this week. Computed nightly from the price ledger since 2026-05-22.
curl 'https://tokenpylon.vercel.app/api/v1/trends?scope=provider:openai&since=2026-08-01'
{"scope":"provider:openai","since":"2026-08-01","until":"2026-09-22","unit":"usd_per_million_tokens","points":[{"day":"2026-08-01","listings":41,"sources_covered":3,"sources_total":3,"median_input":1.25,"median_output":10,"p25_input":0.4,"p75_input":2.5,"p25_output":1.6,"p75_output":10,"min_input":0.05,"min_output":0.4,"cuts":0,"raises":0,"detail":null}]}Pick the cheapest / greenest way to run a workload across hosted-API, hosted-GPU, and self-deploy.
/api/v1/cost/best-deployCheapest end-to-end way to run a workload.
Returns up to ~15 deployment options ranked by cost. Set useLiveCarbon=true to overlay current grid carbon onto self-deploy options.
curl -X POST 'https://tokenpylon.vercel.app/api/v1/cost/best-deploy' -H 'content-type: application/json' \
-d '{"family":"llama-3-1-70b-instruct","inputTokens":1000000,"outputTokens":200000,"useLiveCarbon":true}'{"cheapest_by_cost":{"provider":"fireworks_ai","cost_per_million_tokens":0.100,...},"greenest_by_emissions":{...},"options":[...]}/api/v1/schedule/best-windowTime-shift scheduler: when + where to run for lowest carbon.
Given a workload + deadline window in hours, recommends the (city, hour) tuple that minimizes emissions using historical hour-of-day carbon patterns.
curl -X POST 'https://tokenpylon.vercel.app/api/v1/schedule/best-window' -H 'content-type: application/json' \
-d '{"family":"llama-3-1-70b-instruct","inputTokens":1000000,"outputTokens":200000,"deadlineHours":24}'{"recommended":{"city_code":"us-quincy","suggested_hour_utc":6,"estimated_emissions_kg":0.11,...},"baseline_if_run_now":{...},"reduction_pct":89}Where the compute lives — provider locations, datacenter announcements, electricity.
/api/v1/provider-locationsGPU provider datacenter locations.
46 locations across Vultr (33) + Modal (13). Joined to electricity_rates.city_code where the city has carbon data.
curl 'https://tokenpylon.vercel.app/api/v1/provider-locations?has_carbon=true'
{"locations":[{"provider":"vultr","city_name":"Frankfurt","city_code":"eu-frankfurt"}]}/api/v1/electricityPer-city electricity rates + carbon intensity + PUE.
19 cities covered. Each row has retail rate, datacenter PUE estimate, and annual carbon intensity from eGRID / Ember.
curl 'https://tokenpylon.vercel.app/api/v1/electricity'
{"cities":[{"city_code":"us-quincy","rate_usd_kwh":0.05,"carbon_intensity_gco2_kwh":286.5}]}The AI infrastructure economy — chips, datacenters, training runs, capex, sanctions.
/api/v1/marketSingle-shot dashboard summary across all 13 datasets.
Counts + top-5 for each table. Use this to discover the other endpoints.
curl 'https://tokenpylon.vercel.app/api/v1/market'
{"capex":{...},"gpu_clusters":{"count":680,...},"training_runs":{...},"bis_entity_list":{"count":3399,...}}/api/v1/market/capexHyperscaler quarterly capex from SEC EDGAR.
5 hyperscalers (MSFT/GOOGL/META/AMZN/AVGO). ~10 years of history. Filters: ?ticker=, ?form=10-Q|10-K, ?since=.
curl 'https://tokenpylon.vercel.app/api/v1/market/capex?ticker=MSFT&since=2024-01-01'
{"rows":[{"ticker":"MSFT","period_end":"2026-03-31","capex_usd":30900000000}]}/api/v1/market/tsmcTSMC monthly net revenue.
Source: SEC EDGAR 6-K filings (TSMC's own site is Cloudflare-blocked). Includes MoM%, YoY%, YTD.
curl 'https://tokenpylon.vercel.app/api/v1/market/tsmc'
{"rows":[{"month":"2026-04-01","revenue_twd_m":410725,"yoy_pct":17.5}]}/api/v1/market/clustersFrontier GPU clusters with lat/lon, MW, owner.
From Epoch AI. ?current=true filters to non-superseded snapshots (avoids triple-counting xAI Colossus Phase 1+2+3).
curl 'https://tokenpylon.vercel.app/api/v1/market/clusters?current=true&sort=h100&limit=10'
{"rows":[{"cluster_name":"xAI Colossus Memphis Phase 3","owner":"xAI","h100_equivalents":275796}]}/api/v1/market/training-runsNotable ML training runs with compute + cost.
From Epoch AI's Notable Models dataset. 1,074 rows back to 1950. Filters: ?org=, ?frontier=true, ?since=.
curl 'https://tokenpylon.vercel.app/api/v1/market/training-runs?frontier=true&sort=cost&limit=5'
{"rows":[{"model_name":"Grok 4","organization":"xAI","training_cost_2023_usd":388000000}]}/api/v1/market/hardwareAI accelerator specs (FP4..FP64, memory, TDP).
175 chips: NVIDIA, AMD, Google TPU, Amazon Trainium, Meta MTIA, Cerebras, Groq, Graphcore, Huawei Ascend.
curl 'https://tokenpylon.vercel.app/api/v1/market/hardware?manufacturer=NVIDIA&sort=date'
{"rows":[{"hardware_name":"NVIDIA GB300","release_date":"2025-08-22","fp8_flops":5e15,"tdp_watts":1400}]}/api/v1/market/companiesAI lab financials (revenue, valuation, staff).
9 AI labs from Epoch AI. Sort by ?sort=valuation|revenue|staff|funding.
curl 'https://tokenpylon.vercel.app/api/v1/market/companies?sort=revenue'
{"rows":[{"company_name":"Anthropic","valuation_usd_latest":380000000000,"annual_revenue_usd_latest":30000000000}]}/api/v1/market/datacentersReal-time datacenter buildout announcement firehose.
DCK + DCD + Bisnow RSS feeds, LLM-extracted (operator, city, MW, online_date). Default returns only extracted buildouts; use ?status=all.
curl 'https://tokenpylon.vercel.app/api/v1/market/datacenters?operator=Microsoft'
{"rows":[{"title":"...","operator":"...","city_name":"...","power_mw":120}]}/api/v1/market/ppasHyperscaler renewable PPA + nuclear restart firehose.
MSFT / Google / Meta / Amazon press feeds, LLM-classified for energy deals.
curl 'https://tokenpylon.vercel.app/api/v1/market/ppas?buyer=Microsoft'
{"rows":[{"buyer":"Microsoft","generation_type":"nuclear-restart","capacity_mw":835}]}/api/v1/market/top500Top500 supercomputer rankings biannually since 1993.
Last 10 years by default. Filters: ?list=2025-11-01, ?country=, ?gpu=H100.
curl 'https://tokenpylon.vercel.app/api/v1/market/top500'
{"rows":[{"rank":1,"system_name":"El Capitan","rmax_pflops":1809,"gpu_chip":"MI300A"}]}/api/v1/market/export-controlsUS Consolidated Screening List (Entity List + 8 other lists).
25k+ entities. Filters: ?list=entity|meu|cmic|sdn|all, ?country=CN, ?since=2024-01-01.
curl 'https://tokenpylon.vercel.app/api/v1/market/export-controls?list=entity&country=CN&since=2024-01-01'
{"rows":[{"name":"...","start_date":"2025-10-08","addresses":"Beijing..."}]}/api/v1/market/interconnectionUS grid interconnection queues from PJM/MISO/NYISO/CAISO/ERCOT.
19,383 projects. Filters: ?iso=, ?state=, ?fuel=, ?status=, ?min_mw=.
curl 'https://tokenpylon.vercel.app/api/v1/market/interconnection?iso=ERCOT&min_mw=1000&fuel=GAS'
{"rows":[{"project_id":"30INR0071","capacity_mw":1660,"county":"Franklin","fuel_type":"GAS"}]}Audit what you spend, find the same model cheaper, consider other models with an honest requires_eval flag, deploy a routing policy, then verify the saving on cost per successful task. Every number is a list price with an observation time; nothing here knows your negotiated rate.
/api/v1/cost/quoteItemised workload quote: uncached input, cache reads/writes, output, batch rates, long-context tiers.
Body: { model, provider?, inputTokens, outputTokens, cachedTokens?, cacheWriteTokens?, batch?, calls? }. inputTokens excludes cached tokens. Long-context tiers apply automatically when the prompt exceeds a published threshold. Returns lines with the ledger row each rate came from, unknowns (every assumption), complete=false when a rate was missing, price_observed_at and a 24h quote_expires_at.
curl -X POST 'https://tokenpylon.vercel.app/api/v1/cost/quote' -H 'content-type: application/json' -d '{"model":"claude-sonnet-4-5","inputTokens":50000,"cachedTokens":200000,"outputTokens":4000,"batch":true,"calls":1000}'{"model_key":"claude-sonnet-4-5","provider":"anthropic","per_call":0.1,"total":100,"complete":true,"applied":{"batch":true,"context_tier":"above-200k","caching":true},"unknowns":[],"lines":[{"kind":"input","tokens":50000,"rate_per_million":3,"cost":0.15,"from":"claude-sonnet-4-5:batch"}],"provenance":{"price_observed_at":"...","quote_expires_at":"..."}}/api/v1/cost/auditSpend audit: your usage repriced at today's rates, plus the cheapest other host selling the same model.
Body: { lines: [{ model, provider?, inputTokens, outputTokens, cachedTokens?, calls?, actualCostUsd?, label? }], period?, allowedProviders? }. Same-model-other-host is the one saving that needs no quality judgement. Lines that can't be fully priced come back under unresolved, never zeroed. Batch options are shown where a batch rate is published.
curl -X POST 'https://tokenpylon.vercel.app/api/v1/cost/audit' -H 'content-type: application/json' -d '{"period":"2026-08","lines":[{"model":"meta-llama/llama-3.3-70b-instruct","provider":"together_ai","inputTokens":8000,"outputTokens":1200,"calls":50000}]}'{"totals":{"baseline":492,"cheapest_same_model":60,"savings":432,"savings_pct":87.8},"lines":[{"model_key":"...","provider":"together_ai","cheapest_same_model":{"provider":"deepinfra","total":60,"savings_pct":87.8}}],"unresolved":[]}/api/v1/decide/alternativesCheaper models for the same task, within a quality tolerance of the model you use. Every row requires_eval.
Body: { model, task?, qualityTolerance?, allowedProviders?, minTokensPerSec?, maxTtftMs?, openWeightsOnly?, requires?, limit? }. The bar is YOUR model's benchmark score on the task metric, not a fixed floor: LiveBench coding for coding, LiveBench average for chat and cheap-batch, SimpleBench for reasoning. A model LiveBench has not scored is compared on Aider polyglot (frozen since October 2025) against other Aider results only; baseline.metric names the one used. Capability requirements (tools, json, vision, contextTokens) are echoed back as not_enforced: the catalog carries no capability data yet.
curl -X POST 'https://tokenpylon.vercel.app/api/v1/decide/alternatives' -H 'content-type: application/json' -d '{"model":"gpt-4.1","task":"coding","qualityTolerance":0.1}'{"baseline":{"model_key":"gpt-4.1","score":0.52,"cost_per_run_usd":9.9},"alternatives":[{"model_key":"deepseek/deepseek-chat","score":0.50,"cost_per_run_usd":0.88,"cost_vs_baseline_pct":-91,"requires_eval":true,"evidence":"aider-polyglot 50.0% vs your 52.0%; ..."}]}/api/v1/cost/verifyDid the switch save money? Cost per successful task, baseline vs candidate, with failures, retries and fallbacks counted.
Body: { baseline: [lines], candidate: [lines] } where a line is usage plus calls, successes/failures, retries, fallbacks and optionally actualCostUsd. A cheaper model that fails more is not a saving. Stateless: nothing is stored.
curl -X POST 'https://tokenpylon.vercel.app/api/v1/cost/verify' -H 'content-type: application/json' -d '{"baseline":[{"model":"gpt-4.1","inputTokens":8000,"outputTokens":1000,"calls":1000,"failures":20}],"candidate":[{"model":"deepseek/deepseek-chat","inputTokens":8000,"outputTokens":1100,"calls":1000,"failures":45,"retries":30}]}'{"baseline":{"cost_per_successful_task":0.0244},"candidate":{"cost_per_successful_task":0.0031},"delta":{"pct":-87.3,"verdict":"saving"},"notes":["candidate success rate is 2.5 points lower; ..."]}The routing brain as a remote MCP server, plus a LiteLLM router config priced from live data. Your gateway does the routing; TokenPylon keeps its price table honest.
/api/mcpRemote MCP server (Streamable HTTP). Tools: audit_spend, quote_job, cheaper_alternatives, verify_savings, decide, best_provider, estimate_cost, compare_cost, project_spend, families, price_changes, media_prices, litellm_router_config, watch_model.
Attach it once and your agent can pick a model for a job, find the cheapest provider under speed/latency constraints, price the work, and export a LiteLLM config. Stateless, no auth, same rate limit as the REST API. Claude Code: claude mcp add --transport http tokenpylon https://tokenpylon.vercel.app/api/mcp
curl -X POST 'https://tokenpylon.vercel.app/api/mcp' -H 'content-type: application/json' -H 'accept: application/json, text/event-stream' -d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}'{"jsonrpc":"2.0","id":1,"result":{"tools":[{"name":"decide",...},{"name":"best_provider",...}]}}/api/v1/route/litellmLiteLLM Router model_list with cost-based routing, priced from live data.
?family=a,b names families; ?task=coding&priority=balanced&top=5 lets /decide pick them; ?providers= restricts deployments; ?format=yaml returns a drop-in config.yaml. Deployments carry input_cost_per_token / output_cost_per_token so LiteLLM's cost-based-routing and spend tracking see real prices. Add your own api_key / api_base per provider. Regenerate daily.
curl 'https://tokenpylon.vercel.app/api/v1/route/litellm?family=llama-3-3-70b-instruct&format=yaml'
model_list:
- model_name: llama-3-3-70b-instruct
litellm_params:
model: deepinfra/meta-llama/Llama-3.3-70B-Instruct-Turbo
input_cost_per_token: 1.0e-7
output_cost_per_token: 3.2e-7
router_settings:
routing_strategy: cost-based-routingA small signed binary you install once. Your apps point their base URL at it (or you run them through it); every call is forwarded unchanged and its usage metadata recorded per attempt — model asked for and served, billing provider and gateway, tokens by kind, provider-reported cost, timing, outcome — never messages, never keys. A local page shows your usage priced at today's list rates and the one switch worth making; the catalog gets prices verified by real workloads.
tokenpylon upInstalls the collector as a login service on 127.0.0.1:4141 and opens the local usage page: totals, tokens and cost per token by source (proxy, LiteLLM, Claude Code, Codex), spend at list, and what other contributors pay on the same models. `tokenpylon mcp` is an MCP server any agent can query (live sessions and context, quota meters, usage, priced breakdown); `tokenpylon hook context` warns a harness past 100k tokens of context. Any client can tag its calls with X-Tokenpylon-Session, X-Tokenpylon-Agent, X-Tokenpylon-Project and X-Tokenpylon-Harness headers (stripped before forwarding) to group on the page by session and agent; without them the User-Agent names the harness and calls less than 15 minutes apart form a session. Claude Code and Codex usage is read from their own history files automatically; for everything else, `tokenpylon connect shell` (or claude-code, codex) or `tokenpylon run -- <your command>`.
Base-URL adapters: /openai/v1, /anthropic, /gemini, /openrouter/v1, /deepseek/v1, /mistral, /groq/v1, ... and /proxy/<api-host> for anything OpenAI-shaped. Streaming passes through unbuffered; usage is read from the provider's final event; a client that disconnects tears the upstream call down and the attempt is recorded as cancelled. `tokenpylon import litellm <spend-logs.json>` and `tokenpylon import openrouter <activity.csv>` bring in history; the LiteLLM hook (pip install tokenpylon-litellm) records every attempt the proxy makes. Sharing call metadata with tokenpylon.com is on by default and disclosed at first run: `tokenpylon sharing off` or TOKENPYLON_SHARE=0 turns it off, and the local page works either way. `tokenpylon sharing show` prints the exact fields. Install: npm i -g tokenpylon, brew install tokenpylon/tap/tokenpylon, or the container image for a sidecar.
npm i -g tokenpylon && tokenpylon up && tokenpylon connect shell
TokenPylon sends call metadata to tokenpylon.com by default; never messages or API keys. We use it for your audit and shared price evidence, and retain it indefinitely. Turn sharing off: tokenpylon sharing off — or set TOKENPYLON_SHARE=0. collector installed (~/Library/LaunchAgents/com.tokenpylon.collector.plist) dashboard: http://127.0.0.1:4141/
/api/v1/telemetry/auditWhat the local page calls: your usage totals per (provider, model, host) priced at list, the cheapest same-model host for that exact token mix, the one opportunity worth acting on, and the catalog keys to watch.
Body: { period?, lines: [{ provider, model, served_host, source?, calls, input_tokens, cached_tokens, cache_write_tokens, cache_write_1h_tokens?, output_tokens, reported_cost_usd?, reported_calls? }] }. cache_write_1h_tokens is the part of the writes at the one-hour cache TTL, priced at its own tier. Stateless and unauthenticated: nothing is stored, so it works with sharing off. Only a provider-reported charge counts as a bill; LiteLLM's cost-map figure is calculated, not billed, and never verifies a price. An opportunity needs at least $1 and 5% to be named.
curl -X POST 'https://tokenpylon.vercel.app/api/v1/telemetry/audit' -H 'content-type: application/json' -d '{"lines":[{"provider":"together_ai","model":"meta-llama/Llama-3.3-70B-Instruct-Turbo","served_host":"together_ai","calls":5000,"input_tokens":5000000,"cached_tokens":0,"cache_write_tokens":0,"output_tokens":1000000}]}'{"totals":{"baseline":5.28,"cheapest_same_model":0.82,"savings":4.46,"savings_pct":84.5,"reported":0},"opportunity":{"kind":"switch_host","headline":"$4.46 (84%) on meta-llama/Llama-3.3-70B-Instruct-Turbo by calling it at deepinfra",...},"watch":["meta-llama/Llama-3.3-70B-Instruct-Turbo"]}/api/v1/telemetry/aggregatesThe return feed: what contributed real traffic says about a model's hosts, published only above 5 installs, 1,000 calls, and no install over half.
?model=&provider=&days=30. Per (provider, model, host, cache band, context band): what providers actually charged per million tokens where they report cost, the catalog's quote for the same usage, the gap between them, error rate, and observed time-to-first-token and latency. Raw events are POSTed by the collector to /api/v1/telemetry/usage using a TokenPylon account API key (Authorization: Bearer). Raw records are isolated by account; public cohorts remain shared. Configure the recorder with tokenpylon auth set --stdin. Uploads stay queued until a key is configured. Events use a strict field allowlist (schema v2: per attempt, with logical call id, requested vs returned model, gateway vs billing provider, and an amount basis on every dollar figure) and kept indefinitely; cohorts are rolled up nightly.
curl 'https://tokenpylon.vercel.app/api/v1/telemetry/aggregates?model=deepseek-v4-pro'
{"count":2,"aggregates":[{"provider":"deepseek","model":"deepseek-v4-pro","served_host":"deepseek","calls":4120,"installs":9,"reported_usd_per_million":1.48,"quoted_usd_per_million":1.51,"reported_vs_quoted_pct":-2.0,"error_rate":0.004,"p50_ttft_ms":610}]}A scraped price is a claim. Every night TokenPylon makes one tiny real call against a rolling slice of listings with its own keys, reads the bill back where the provider allows it, and records billed vs quoted. Quotes carry the result.
/api/v1/health/verificationCoverage of the verification job: listings verified by method, recent mismatches, spend to date, the caps in force.
Methods: response_cost (the provider returned dollars, e.g. OpenRouter), balance_delta (a prepaid balance moved by exactly the call), usage_only (token accounting checked, bill not observable in-band), unreachable. A quote's `verification` field carries the latest row for that listing, or null when never tested. Spend is capped per run and per month and every call is one ledger row.
curl 'https://tokenpylon.vercel.app/api/v1/health/verification'
{"enabled":true,"coverage":{"listings_verified":412,"reachable":398,"by_method":{"response_cost":260,"balance_delta":90,"usage_only":48,"unreachable":14}},"spend":{"month_to_date_usd":0.41},"mismatches":[{"model_key":"...","provider":"...","delta_pct":6.2}]}Internal health endpoints exposed publicly so anyone can audit data freshness.
/api/v1/health/freshnessDays-since-last-success per data source.
Surfaces stale scrapers before they affect any chart.
curl 'https://tokenpylon.vercel.app/api/v1/health/freshness'
{"summary":{"fresh":7,"stale":0,"dead":0},"sources":[{"source":"litellm","status":"fresh","days_since_last_success":0.9}]}/api/v1/health/coveragePer-source row counts and last-update times.
How many rows each scraper has loaded and when it last successfully ran.
curl 'https://tokenpylon.vercel.app/api/v1/health/coverage'
{"sources":[{"source":"epoch-gpu-clusters","row_count":680,"last_run_at":"..."}]}/api/v1/health/cross-checkDisagreement between independent sources.
Where two sources report the same data point, surface which pairs disagree most. Used to find suspect numbers.
curl 'https://tokenpylon.vercel.app/api/v1/health/cross-check'
{"rows":[]}