ModelPriceWatch.com
Last scan 2026-08-14 Models tracked 204 Providers 31 Cheapest paid Granite 4.0 H Micro $0.017/Mtok in Every price links to its source

Granite 4.0 H Micro

by IBM

Current budget open weights cheap tier
Availability: This row is IBM's HYBRID 3B (4 attention / 36 Mamba2 layers), not the dense 3B Granite 4.0 Micro that IBM ships alongside it. The two are separate models on IBM's own model cards, and only the hybrid has a per-token price: Cloudflare Workers AI serves @cf/ibm-granite/granite-4.0-h-micro, while Artificial Analysis states the dense Granite 4.0 Micro "is not currently available through any API providers we benchmark". Until 2026-08-14 we published this price under the dense model's name and benchmarks; OpenRouter titles the same granite-4.0-h-micro slug "IBM: Granite 4.0 Micro", which is how the two were conflated. This is Cloudflare's price for hosting IBM's model, not a price from IBM: IBM's watsonx.ai rate card lists granite-4h-micro as "Not available" for per-million-token pay-as-you-go while pricing its siblings on the same table (granite-4h-small is USD 0.0636 in / USD 0.265 out), and sells it only as hourly on-demand GPU hosting, which is not comparable per token. The output rate is $0.112 to three decimals on Cloudflare's Workers AI pricing table; the model page rounds it to $0.11.

Today's price · per 1M tokens

Input

$0.017

per 1M tokens

Output

$0.112

per 1M tokens

Blended

$0.041

blended $/1M — 3:1 weighted input:output

Source: official IBM pricing · read Aug 14, 2026 MODELPRICEWATCH.COM · 2026-08-14

Price receipt

No confirmed receipt. We hold dated captures of IBM's pricing page, but none of them tied Granite 4.0 H Micro to a published price on that page, so we do not claim this number is sourced. Check the provider's page directly before relying on it.

Overview

Hybrid 3B Granite 4 model, sold per token by Cloudflare Workers AI rather than by IBM. $0.017/$0.112 per 1M tokens.

Run it yourself

Deploy this open model on rented GPUs

Open weights mean you can self-host instead of paying per-token API prices. These platforms let you serve it on demand — often cheaper at scale.

Partner links — we may earn a commission if you sign up, at no cost to you. We only list platforms we'd recommend regardless.

Capabilities

struck through = not supported
Input 1/5
Text ✓ Image Audio Video PDF
Output 1/5
Text ✓ Image Audio Video Embedding
Features 3/9
Prompt caching Reasoning Coding Fast inference ✓ Long context ✓ Open weights ✓ Multimodal Web search Realtime

Benchmark performance

accuracy % · higher is better
MMMLU vendor
55.19%
BFCL vendor
57.56%
HumanEval vendor
81%

Every per-benchmark score we hold for this model is reported by its own vendor, so none is counted toward an accuracy average or any ranking.

Source: IBM model card, IBM model card (5-shot) · updated Aug 14, 2026 · See full rankings →

Specifications

Provider
IBM
Context window
128K tokens
Modality
text
Parameters
Proprietary
Open source
Yes — open weights available
Released
Oct 19, 2025
Status
Current
Last updated
Aug 14, 2026
Tags
open-weightsbudgetfast

Availability verified: Aug 14, 2026 — listed on IBM's own page