Base Models

Discover all the foundation models available on LetzAI

inferencesh

Seedance 2.5 Video Upscale

ByteDance
Aug 2026

Upscale a video to 1080p via Seedance 2.5 reference-to-video. Max 30s. 4K is not available on 2.5.

inferencesh

Seedance 2.5 Enterprise Video Upscale

ByteDance
Aug 2026

Org-scoped Seedance 2.5 Studio video upscale to 1080p. Max 30s. 4K is not available on 2.5.

byteplus

BytePlus Video Enhancer

ByteDance
Aug 2026

Cinema-grade BytePlus MediaKit enhancement (Professional): 30+ algorithms, AI super-res, defect repair, and color. Keeps the original footage (does not regenerate).

xai

Grok 4.6

xAI
Aug 2026

xAI's newest flagship for coding, agentic tool calling, and long-running visual work. 500k context; reasoning effort low/medium/high/xhigh (default high). Knowledge cutoff February 1, 2026.

xai

Grok Imagine Image Quality

xAI
Aug 2026

xAI's quality-focused image model — photorealistic rendering, strong text and logo accuracy, and multi-style versatility from portraits to anime. $0.05/image; 1K and 2K.

inferencesh

Seedance 2.0 Video Upscale

ByteDance
Aug 2026

Upscale a video to 1080p or 4K via Seedance 2.0 reference-to-video. Max 15s.

inferencesh

Seedance 2.0 Enterprise Video Upscale

ByteDance
Aug 2026

Org-scoped Seedance 2.0 Studio video upscale to 1080p or 4K. Max 15s.

xai

Grok Imagine 2.0

xAI
Aug 2026

xAI's newest image model — precise instruction following, sharp typography, and iterative editing with up to 5 reference images. $0.04/image; 1K and 2K.

inferencesh

Seedance 2.5 Enterprise

ByteDance
Aug 2026

ByteDance's professional Seedance 2.5 Studio variant with private asset library support. Automatically uploads reference images to BytePlus virtual portrait library for enhanced character consistency. Up to 30s clips with up to 50 multimodal references (30 images / 10 videos / 10 audio) and synchronized audio. 480p, 720p, and 1080p.

inferencesh

Flux 3 Video

Black Forest Labs
Aug 2026

Black Forest Labs' Flux 3 Video — generate and animate video up to 20s at HD or Full HD with synchronized audio. Supports text-to-video, image-to-video (start frame), and first+last frame keyframes. No Omni / multimodal reference stack.

inferencesh

MiniMax H3

MiniMax
Jul 2026

MiniMax's frontier video model at 2K. Smart-routes text-to-video, image-to-video (optional last frame), and multimodal reference-to-video (reference images, one video, one audio) while keeping subjects consistent.

anthropic

Claude Opus 5

Anthropic
Jul 2026

Anthropic's strongest Opus — step-change over 4.8 for agentic coding, long-horizon work, and professional knowledge tasks. Adaptive thinking; 1M context; near-Fable intelligence at Opus price ($5/$25).

openai

GPT-5.6 Sol

OpenAI
Jul 2026

OpenAI's GPT-5.6 flagship for complex professional work — coding agents, long research, computer use, and multi-tool workflows. Native web search via the Responses API.

byteplus

Seedream 5 Pro

ByteDance
Jul 2026

ByteDance's latest Seedream model with Chain-of-Thought reasoning, enhanced prompt following, reference consistency, and up to 4K resolution.

inferencesh

Seedance 2.5

ByteDance
Jul 2026

ByteDance's next-generation Seedance with up to 30-second native clips, up to 50 multimodal references (30 images / 10 videos / 10 audio), and synchronized audio. Supports text, image, first+last frame, and reference-to-video. 480p, 720p, and 1080p.

fal

GPT Image 2 Upscaler

Black Forest Labs
Jul 2026

AI upscaling via OpenAI GPT Image 2 edit with intelligent detail enhancement up to 4K resolution.

anthropic

Claude Sonnet 5

Anthropic
Jun 2026

Anthropic's latest Sonnet — drop-in upgrade with adaptive thinking, stronger agentic coding, and 1M context. Introductory pricing through Aug 2026.

vertex

Nano Banana 2 Lite

Google
Jun 2026

Google's fastest and cheapest Gemini image model — optimized for high-volume generation where speed and cost matter most. Best for single-prompt text-to-image; not optimized for multiple reference inputs or multi-turn editing.

inferencesh

Gemini Omni Flash

Google
Jun 2026

High-speed multimodal video with native audio — the cheaper, faster Omni-reference alternative.

inferencesh

Nano Banana 2 Upscaler

Google
Jun 2026

Faster AI upscaling using Google Gemini 3.1 Flash with intelligent detail enhancement up to 4K resolution.

inferencesh

Pruna Upscaler

Pruna AI
May 2026

Fast AI upscaling via Pruna P-Image-Upscale with optional detail and realism enhancement up to 128 megapixels.

beeble

Beeble SwitchX

Beeble
May 2026

Beeble's video-to-video model that swaps the background or scene in a clip while preserving the original subject, motion, framing and lighting. Not for character or outfit changes. Driven by an optional reference image plus prompt.

inferencesh

Seedance 2.0 Enterprise

ByteDance
May 2026

ByteDance's professional Seedance 2.0 Studio variant with private asset library support. Supports up to 4K (10-bit color), with text-to-video, image-to-video, and multimodal reference-to-video (up to 9 reference images, 3 reference videos, 3 reference audios) and synchronized audio.

openai

GPT-5.5

OpenAI
Apr 2026

OpenAI's frontier model for complex reasoning, coding, and agentic work. Native web search via the Responses API.

fal

GPT Image 2

OpenAI
Apr 2026

OpenAI's GPT Image 2 via FAL — exceptional typography, fine-detail rendering, and precise mask-based editing up to 4K.

byteplus

Seedance 2.0

ByteDance
Apr 2026

ByteDance's most advanced video model with cinematic output, native audio, real-world physics, and director-level camera control. Supports up to 4K (10-bit color). Supports text, image, audio, and video reference inputs (up to 9 reference images, 3 reference videos, 3 reference audios) with first+last frame control.

inferencesh

WAN 2.7 Image Pro

Alibaba
Apr 2026

Alibaba's professional image generation model supporting text-to-image, image editing, and multi-reference generation with up to 4K high-definition output. Supports thinking mode for improved quality.

vertex

Nano Banana 2

Google
Feb 2026

Google's Gemini 3.1 Flash image generation model — fast, cost-effective generation and editing with up to 4K resolution, advanced text rendering, and Google Search grounding.

inferencesh

Nano Banana 2

Google
Feb 2026

Google Gemini 3.1 Flash via InferenceSH — fast, cost-effective image generation and editing with up to 4K resolution and Google Search grounding.

fal

Kling O3 Pro Edit

Kuaishou
Feb 2026

Kuaishou's unified multimodal video editor with 7-in-1 capabilities including object removal, background swapping, style changes, and character consistency.

fal

Kling V3

Kuaishou
Feb 2026

Kuaishou's flagship video model with native multilingual audio, up to 15 seconds, multi-shot storytelling, and reference image support.

byteplus

Seedream 4.5

ByteDance
Dec 2025

ByteDance's image model with multi-image consistency, excellent text rendering, and up to 4K resolution. Maintains character identity across multiple images.

fal

Kling V2.6

Kuaishou
Dec 2025

Kuaishou's video model with simultaneous audio-visual generation, creating videos with voiceovers and sound effects in a single pass.

fal

Flux 2

Black Forest Labs
Nov 2025

Black Forest Labs' 32B parameter model with multi-reference image support, 4MP editing, improved text rendering, and enhanced photorealism.

vertex

Nano Banana Pro

Google
Nov 2025

Google's state-of-the-art image generation model with improved text rendering, multi-turn editing, and professional-grade controls over lighting, camera, and composition.

vertex

Nano Banana Pro

Google
Nov 2025

Google Gemini 3 Pro via InferenceSH with improved retries and per-organization API key support.

inferencesh

Nano Banana Pro Upscaler

Google
Nov 2025

AI-powered upscaling using Google Gemini 3 Pro with intelligent detail enhancement up to 4K resolution.

vertex

Veo 3.1 Extend Video

Google
Oct 2025

Google's video extension model capable of extending existing videos up to 148 seconds with maintained consistency and style.

vertex

Google VEO3.1

Google
Oct 2025

Google's updated video model via Vertex AI with richer native audio, improved dialogue sync, and enhanced cinematic style.

fal

WAN 2.5

Alibaba
Sep 2025

Alibaba's open-source video model with multilingual support, synchronized audio, and up to 10 seconds at 1080p resolution.

How we choose models

Tested by creative professionals

We use our own suite of benchmarks to identify the best AI models from leading providers to serve you only the best and most useful creative tools at all times.

API Access

Integrate LetzAI into your apps, workflows, and agents with our REST API.

1
Get your API key

Subscribe to any paid plan, then find your key on the subscription page.

2
Generate an image

Send a POST request to create an image with any model.

POST https://api.letz.ai/images
3
Retrieve your result

Poll the image endpoint with the returned ID until it's ready.

GET https://api.letz.ai/images/:id

Our Inference Providers

We partner with leading AI companies to bring you the best models.

Model availability and pricing may change. Check the pricing page for current rates.