OpenAIOpenAI·text

OpenAI GPT OSS 120B

ReasoningWeb searchFunction calling
Quick reference
OpenAI GPT OSS 120B — TLDR
  • - 🆕 OpenAI's open-weight Mixture-of-Experts model under Apache 2.0 license.
  • - 📏 117B total parameters, 5.1B active per token.
  • - ⚡ Fits on a single 80GB H100 or AMD MI300X GPU.
  • - 🧠 Configurable reasoning effort: low, medium, or high.
  • - 🔧 Native function calling, browsing, Python execution, structured outputs.
  • - 👁️ Full chain-of-thought access for debugging and inspection.
  • - 📚 128K context window; fine-tunable on a single H100 node.
  • - 🔒 Native MXFP4 quantization on MoE weights for efficient deployment.
💰 Best price on AntSeed
$0.0064 / $0.02791%
per 1M · cheapest in / out
📏 Context
128K tokens
🐜 Sellers
6
advertising on AntSeed
Provider

OpenAI is an American artificial intelligence research organization headquartered in San Francisco, structured as both a for-profit public benefit corporation and a nonprofit foundation. The lab developed the GPT family of large language models, the DALL-E image generation…

Explore 20 more models by OpenAI
About this model

OpenAI GPT OSS 120B is the larger of OpenAI's two open-weight gpt-oss models, designed for production-grade reasoning, agentic workflows, and general-purpose use under the permissive Apache 2.0 license. Architecturally it is a Transformer Mixture-of-Experts model with 36 layers, 128 experts per layer (4 active per token), and roughly 117B total parameters of which about 5.1B are active per forward pass. Native MXFP4 quantization of the MoE weights lets it run on a single 80GB GPU such as an NVIDIA H100 or AMD MI300X.

Within the family, it sits above the smaller [[sibling:e2ee-gpt-oss-20b-p|GPT OSS 20B]], which carries roughly 21B total and 3.6B active parameters and is targeted at lower-latency or local deployments that fit within 16GB of memory. The 120B variant trades that footprint for higher reasoning capacity and can itself be fine-tuned on a single H100 node.

Developers can dial reasoning effort across three levels with a single line in the system prompt, and gain full access to the model's chain-of-thought, which OpenAI notes is intended for debugging rather than end-user display. Agentic features include native function calling, web browsing, Python code execution, and structured outputs, with the model able to chain together many sequential browsing calls.

OpenAI evaluated gpt-oss-120b against its own reasoning models including o3, o3-mini, and o4-mini across coding, competition math, health, and agentic tool-use benchmarks at the high reasoning setting.

View source on GitHub ↗View model card on HuggingFace ↗
Sources
developers.openai.comgpt-oss-120b Model | OpenAI API· developers.openai.comopenai.comIntroducing gpt-oss | OpenAI· openai.comdocs.api.nvidia.comopenai / gpt-oss-120b· docs.api.nvidia.comhuggingface.coopenai/gpt-oss-120b · Hugging Face· huggingface.co

This About section is AI-generated from public sources via VeniceStats + Venice inference, with no human editing. It may contain inaccuracies.

Sellers serving OpenAI GPT OSS 120B (6)
SellerReputationInput $/MCached $/MOutput $/MCategoriesAPI

"Best price" and the seller table are live AntSeed catalog data (advertised $/1M tokens, not settled amounts). Reputation = on-chain trust (0-100). Model knowledge (TLDR, provider, About) via the VeniceStats enrichment layer. Advertised catalog, not the model used in any specific purchase.