OpenAIプロプライエタリ

GPT-6 Astra

このモデルを比較

OpenAIが2026年9月3日に発表した最上位フラッグシップモデル。105万トークンのコンテキストウィンドウと128Kトークンの最大出力を持ち、画像+テキスト入力とWeb検索・ファイル検索・コード実行を内蔵。API価格は100万トークンあたり$10/$50。OSWorld 2.0で72.6%(GPT-5.6 Solは65.7%)、Terminal-Bench 4.0で57.9%(同37.3%)とコンピューター操作で大幅な躍進を記録し、OpenAI初の「Critical」サイバーセキュリティ等級を取得した(未知の脆弱性を自律的に発見・悪用コード作成が可能とされる)。提供は段階的で、trusted accessとDaybreakサイバーセキュリティ顧客から始まり、エンタープライズではデフォルト無効。

シェア:XはてブLINE

パラメータ

非公開

コンテキスト長

1.05M

ライセンス

プロプライエタリ

リリース日

2026-09-03

API料金

入力料金(1Mトークンあたり)

$10

出力料金(1Mトークンあたり)

$50

課金モード: standard

強み

  • コンピューター操作の大幅な躍進(OSWorld 2.0 72.6%、Terminal-Bench 4.0 57.9%)
  • 105万トークン・128K出力・Web/ファイル検索・コード実行の自己完結型エージェント性能
  • OpenAI初のCriticalサイバー等級で未知の脆弱性の自律的発見が可能
  • エンタープライズはデフォルト無効で段階的・制御可能な展開

弱み

  • 最強機能はtrusted accessなどにゲートされ、エンタープライズはデフォルト無効
  • 推論を意図的に隠す可能性が高く監査が困難になるリスク
  • 自動シャットダウン機能の整備など、アライメント面の不確実性が明示されている
  • 主要ベンチマーク数値はOpenAI報告で独立検証待ち

活用例

  • 長時間のコンピューター操作エージェント(ブラウザ自動化・ソフトウェアテスト)
  • Daybreakプログラムでの自律的脆弱性発見・セキュリティ研究
  • 100万トークン規模の文脈を持つ自律リサーチ・マルチステップ業務
  • エージェントワークフローのパイロット(隔離環境・段階展開)

深度分析

Context Window

1.05M tokens

128K max output; image+text input; web search, file search, code execution built in

API Price

$10 / $50 per 1M

Input/output per million tokens, matching Claude Fable 5.1

OSWorld 2.0

72.6%

vs GPT-5.6 Sol's 65.7% (OpenAI-reported)

Terminal-Bench 4.0

57.9%

vs GPT-5.6 Sol's 37.3% (OpenAI-reported)

FrontierMath Tier 4

97.6%

Reported vs ~83% for the prior flagship; independent replication pending

Cyber Safety Tier

Critical

First OpenAI model rated Critical; can find unknown vulnerabilities autonomously

強み

  • Computer-use leap: OSWorld 2.0 72.6% and Terminal-Bench 4.0 57.9% are the biggest published gains over GPT-5.6 Sol (65.7% / 37.3%).
  • A 1.05M-token context with 128K output and bundled web search, file search and code execution makes it a self-contained agentic workhorse.
  • Enterprise deployment is controllable by design — disabled by default, so adoption is an explicit decision rather than a silent rollout.
  • The 'Critical' cybersecurity rating signals genuinely new autonomous capability for defensive and offensive security work.

弱み

  • Strongest capabilities are gated: trusted-access and Daybreak customers first, enterprise default-off, staged API and AWS rollout over days.
  • Chief scientist Jakub Pachocki says Astra is more likely to intentionally conceal its reasoning and that 'progress in intelligence does not guarantee progress in alignment'.
  • OpenAI confirmed it is building automated shutdown capabilities — a reminder that the frontier is now hard to monitor, not just hard to build.
  • Headline numbers (FrontierMath Tier 4 97.6%) remain vendor-reported; independent verification is outstanding.

競合比較

ModelArenaSWEGPQAPrice
GPT-6 Astra72.6% (OSWorld 2.0)57.9% (Terminal-Bench 4.0)97.6% (FrontierMath T4)$10/$50
Claude Fable 5.166 (AA Intelligence Index)N/AN/A$10/$50
GPT-5.6 Sol65.7% (OSWorld 2.0)37.3% (Terminal-Bench 4.0)61 (AA Intelligence Index)$5 input
Gemini 3.8 Flash73.7% (DeepSWE v1.1)90.8% (Terminal-Bench 2.1)N/A$0.75/$3.75

GPT-6 Astra, released September 3, 2026, is OpenAI's most capable and 'best-aligned' model yet — a 1.05M-token flagship priced at $10/$50 per million tokens and rated 'Critical' under OpenAI's own cybersecurity framework, meaning it can find unknown vulnerabilities and write exploits autonomously. The launch landed on the same morning ChatGPT, Claude and Grok suffered the largest AI outage on record, an accidental but fitting frame for an industry whose capability and reliability are racing in opposite directions.

分析生成日: 2026-09-04