GPT Image 2
GPT Image 2 是我们最先进的图像生成模型,用于快速、高质量的图像生成和编辑。它支持灵活的图像尺寸和高保真图像输入。
Access 500+ text, image, video, and audio models with usage-based billing and direct contract pricing. Open any card to view integration docs, SDK examples, and pricing details.
50 models available
GPT Image 2 是我们最先进的图像生成模型,用于快速、高质量的图像生成和编辑。它支持灵活的图像尺寸和高保真图像输入。
Claude Fable 5.1 是 Anthropic 的模型。 模型族:claude-fable。 知识截止:2026-06。 模态:text、image、pdf → text。 上下文长度:1,000,000 tokens。
Claude Opus 5 是 Anthropic 的模型。 模型族:claude-opus。 知识截止:2026-05。 模态:text、image、pdf → text。 上下文长度:1,000,000 tokens。
GPT-6 Astra is OpenAI's flagship frontier model, released September 3, 2026 and described by OpenAI as its most capable and most aligned model, built for the hardest end-to-end work in coding, computer use, browsing, research and professional tasks. On the API it offers a 1,050,000-token context window with up to 128,000 output tokens, a knowledge cutoff of April 30, 2026, text and image input with text output, and support for function/tool calling, built-in tools (web search, file search, computer use), Structured Outputs and streaming; reasoning effort supports low, medium, high, xhigh and max (none is not supported, and custom temperature/top_p and logprob are unavailable). Standard API pricing is $10 per 1M input tokens and $50 per 1M output tokens, with cached input at $1.00, cache writes at $12.50, and a long-context tier above 272K input tokens at $20/$75 ($2.00 cached, $25.00 cache writes); Batch and Flex are priced at 50% of Standard rates and Fast mode at 2x. Access is being rolled out gradually, with parts of it gated behind OpenAI's Trusted Access / Daybreak program as the first model to reach the Critical cybersecurity capability level under its Preparedness Framework.
Claude Opus 4.8 是 Anthropic 的模型。 模型族:claude-opus。 模态:text、image、pdf → text。 上下文长度:1,000,000 tokens。
Claude Sonnet 4.6 是 Anthropic 的模型。 模型族:claude-sonnet。 知识截止:2025-08-31。 模态:text、image、pdf → text。 上下文长度:1,000,000 tokens。
Claude Opus 4.6 是 Anthropic 的模型。 模型族:claude-opus。 知识截止:2025-05-31。 模态:text、image、pdf → text。 上下文长度:1,000,000 tokens。
GPT-5.6 Sol 是 OpenAI 的模型。 模型族:gpt。 知识截止:2026-02-16。 模态:text、image、pdf → text。 上下文长度:1,050,000 tokens。
Claude Opus 4.7 是 Anthropic 的模型。 模型族:claude-opus。 知识截止:2026-01-31。 模态:text、image、pdf → text。 上下文长度:1,000,000 tokens。
GPT-5.6 Terra 是 OpenAI 的模型。 模型族:gpt-mini。 知识截止:2026-02-16。 模态:text、image、pdf → text。 上下文长度:1,050,000 tokens。
GPT-5.6 Luna is a model by OpenAI. Model family: gpt-nano. Knowledge cutoff: 2026-02-16. Modalities: text, image, pdf → text. Context length: 1,050,000 tokens.
GPT-5.5 是 OpenAI 的模型。 模型族:gpt。 知识截止:2025-12-01。 模态:text、image、pdf → text。 上下文长度:1,050,000 tokens。
Qwen3.7 Plus 是 Alibaba Qwen 的模型。 模型族:qwen。 知识截止:2025-04。 模态:text、image、video → text。 上下文长度:1,000,000 tokens。
Qwen3.7 Max 是 Alibaba Qwen 的模型。 模型族:qwen。 模态:text → text。 上下文长度:1,000,000 tokens。
Qwen3.6 Plus 是 Alibaba Qwen 的模型。 模型族:qwen。 知识截止:2025-04。 模态:text、image、video → text。 上下文长度:1,000,000 tokens。
Qwen3.6 Max Preview 是 Alibaba Qwen 的模型。 模型族:qwen。 知识截止:2025-04。 模态:text → text。 上下文长度:262,144 tokens。
Qwen3.6 Flash 是 Alibaba Qwen 的模型。 模型族:qwen3.6。 模态:text、image、video → text。 上下文长度:1,000,000 tokens。
Qwen3.5 Plus 是 Alibaba Qwen 的模型。 模型族:qwen。 知识截止:2025-04。 模态:text、image、video → text。 上下文长度:1,000,000 tokens。
Kimi K3 是 Moonshot AI / Kimi 的模型。 模型族:kimi-k3。 模态:text、image、video → text。 上下文长度:1,048,576 tokens。
HappyHorse-1.1-I2V supports image-to-video generation, further enhancing visual quality, dynamic performance, and cross-clip consistency. The model can more accurately understand input images and continue the creative intent, bringing significant improvements in character skin texture, cross-clip ID consistency, motion fluidity, text rendering stability, and audio-visual synchronization, outputting high-quality videos that are more realistic, natural, rich in detail, and highly consistent.
HappyHorse-1.1-T2V supports text-to-video generation, further enhancing text semantic understanding, camera scheduling, and dynamic generation performance. The model can more accurately restore creative intent, generating higher-quality videos that are smoother, more natural, richer in detail, and more consistent in terms of character movements, scene atmosphere, visual aesthetics, and physical motion.
Hy3 is a 295B-parameter Mixture-of-Experts language model from Tencent (21B active parameters, 192 experts with top-8 routing, plus a 3.8B MTP layer), developed by the Tencent Hy Team and released under Apache 2.0. Refined on the Hy3 preview release with improved post-training data and scaled RL, it targets reasoning, agentic workflows, coding and long-context tasks, supporting a 256K context window, hybrid fast/slow thinking via a configurable reasoning_effort parameter, function calling, structured output and caching. It is served on Tencent Cloud TokenHub with an OpenAI-compatible API.
HappyHorse-1.0-R2V 支持参考生视频,更加稳定的主体与场景参考,支持最多9张图片参考,能够精准保持创作意图,实现更强表现能力。
HappyHorse-1.0-I2V 是阿里巴巴 ATH 团队推出的 15B 参数统一多模态 Transformer 视频生成模型。支持图生视频(Image-to-Video),以图像为起点,精准保留原图内容,生成自然流畅的动态视频。具备高度还原的动态画面生成能力,能够精准理解文本语义,输出流畅自然、细节丰富的高质量视频。支持 720P/1080P 分辨率。
Grok 4.6 是 xAI 的模型。 模型族:grok。 知识截止:2026-02-01。 模态:text、image、pdf → text。 上下文长度:500,000 tokens。
Cinematic creative generation, ultimate dynamic details. 文生视频,精准理解语义,细节丰富画质流畅。支持1080p原生分辨率,同步音频,多镜头故事讲述,7种语言唇形同步。
HappyHorse-1.0 视频编辑模型,支持对已有视频进行风格变换、局部替换等二次创作。支持截取前15秒编辑,输出分辨率与输入一致,实现从1到N的创意延展。
智谱最新旗舰模型,以极致后训练 Scaling 实现能力跃迁。在与 GLM-5.2 完全相同的基座上,依托数十倍规模的长程任务环境、更丰富多样的环境类型与超长周期的后训练,编程体感较前代提升 50%,并在 Terminal Bench 3.0 等公开基准中位列开源模型第一;同时涌现出强大的网络安全能力。
GLM-5.2 is Z.ai's flagship model built for long-horizon tasks, featuring a usable 1M-token context window and 744B MoE parameters (40B activated). It is optimized for project-scale engineering and agentic coding scenarios, with support for multiple reasoning effort levels, function calling, structured output (JSON mode), and streaming.
Gemini 3.5 Flash 是 Google Gemini 的模型。 模型族:gemini-flash。 知识截止:2025-01。 模态:text、image、video、audio、pdf → text。 上下文长度:1,048,576 tokens。
Gemini 3.1 Pro Preview 是 Google Gemini 的模型。 模型族:gemini-pro。 知识截止:2025-01。 模态:text、image、video、audio、pdf → text。 上下文长度:1,048,576 tokens。
Gemini 3.1 Flash Lite Preview 是 Google Gemini 的模型。 模型族:gemini-flash-lite。 知识截止:2025-01。 模态:text、image、video、audio、pdf → text。 上下文长度:1,048,576 tokens。
Nano Banana Lite is the efficiency expert in the image generation family, providing ultra-low latency and highly cost-effective image generation and editing capabilities. Targeting an end-to-end latency of under 2 seconds, it significantly reduces TPU computing costs, making it ideal for high-concurrency interactive developer use cases and real-time consumer applications. It supports interleaved generation and editing from text to image-and-text, and from image to image-and-text, is optimized for 1K (1024x1024px) resolution, supports 14 aspect ratios, enables rapid multi-turn local editing, maintains high character consistency, and has SynthID + C2PA watermarks enabled by default.
Nano Banana 2 是 Google Gemini 的模型。 模型族:gemini-flash。 知识截止:2025-01。 模态:text、image、pdf → text、image。 上下文长度:65,536 tokens。
Nano Banana 2 provides high-quality image generation and conversational editing at mainstream prices with low latency. It is the highly efficient counterpart to Gemini 3 Pro Image, optimized for speed and high-volume developer use cases. It supports 0.5K, 1K, 2K, and 4K resolutions, adds 1:4, 4:1, 1:8, and 8:1 aspect ratios, and supports image search grounding and chain of thought.
Nano Banana Pro 是 Google Gemini 的模型。 模型族:gemini-pro。 知识截止:2025-01。 模态:text、image → text、image。 上下文长度:65,536 tokens。
Nano Banana Pro is a complex reasoning-driven engine for professional-grade image editing and generation, offering studio-level precision and advanced creative control. Nano Banana Pro is best suited for complex graphic design, high-fidelity product modeling, and factual data visualization that requires accurate text rendering and real-world grounding via Google Search.
Nano Banana 是 Google Gemini 的模型。 模型族:gemini-flash。 知识截止:2025-06。 模态:text、image → text、image。 上下文长度:32,768 tokens。
Gemini 2.5 Flash 是 Google Gemini 的模型。 模型族:gemini-flash。 知识截止:2025-01。 模态:text、image、audio、video、pdf → text。 上下文长度:1,048,576 tokens。
Supports multimodal references, featuring video editing and extension capabilities
字节跳动最新一代多模态视频生成模型,支持文本、图片、视频、音频等参考输入,强化角色一致性、复杂运镜、动作表现和长镜头生成,可用于高质量影视级视频创作。
Gemini 2.5 Pro 是 Google Gemini 的模型。 模型族:gemini-pro。 知识截止:2025-01。 模态:text、image、audio、video、pdf → text。 上下文长度:1,048,576 tokens。
新一代专业级多模态视频创作模型,支持基于图像、视频和音频等多模态参考输入生成视频,同时具备视频编辑和扩展能力。
Seedance-2.0-fast is a new generation multimodal video creation model, inheriting the core features and advantages of Seedance-2.0 with faster speed. It supports generating videos from reference images/videos/audio, video editing, video extension, and first and last frame generation.
Claude Sonnet 4.5 是 Anthropic 的模型。 模型族:claude-sonnet。 知识截止:2025-07-31。 模态:text、image、pdf → text。 上下文长度:1,000,000 tokens。
DeepSeek V4 Pro 是 DeepSeek 的模型。 模型族:deepseek-thinking。 知识截止:2025-05。 模态:text → text。 上下文长度:1,000,000 tokens。
Claude Sonnet 5 是 Anthropic 的模型。 模型族:claude-sonnet。 知识截止:2026-01-31。 模态:text、image、pdf → text。 上下文长度:1,000,000 tokens。
DeepSeek V4 Flash 是 DeepSeek 的模型。 模型族:deepseek-flash。 知识截止:2025-05。 模态:text → text。 上下文长度:1,000,000 tokens。
Claude Haiku 4.5 是 Anthropic 的模型。 模型族:claude-haiku。 知识截止:2025-02-28。 模态:text、image、pdf → text。 上下文长度:200,000 tokens。
Claude Fable 5 is Anthropic's most capable publicly released model to date, built specifically for the most demanding reasoning and long-horizon Agent tasks. It supports a 1M context window and up to 128k output tokens. It features built-in safety classifiers and supports adaptive thinking, code execution, visual understanding, and tool use.