APACHE 2.0 OSS · LINUX FOUNDATION AI & DATA

Flyte 2:
The Durable
AI Runtime

Run workloads for training, serving, and
agents, auto-recover from any failure.

Built for scale, performance, and variable compute environments

Get Started Read the docs
✓ Copied install commands
✕ API rate limit

Run a durable, infra-aware workload

Pick a pipeline, pick a failure. Press run and watch Flyte recover — no manual intervention.
Use case:
Break it with:
idle
batch_inference.py
Console
main ▸Ready
Run: r-7f3a2c ⧉  ·  Task: batch_inference.main ⧉
· main 7 15.3m
0m3m6m9m12m15m
API rate limit — A provider returns 429 mid-shard. Flyte catches it, backs off exponentially per the RetryStrategy, and retries the one failed task — completed shards are never rerun.

Write workflows in pure

No DSL, no YAML hell. A TaskEnvironment declares images, resources, secrets and retry policy; tasks are plain async functions. Branch, loop and fan out with ordinary control flow — the runtime records IO for every step so a crash resumes instead of restarting.

flyteorg/flyte-sdk ↗ pip install flyte
→Type-checked I/O between tasks — language-native types, files, DataFrames, dataclasses
→Native agents: wrap tools as tasks, every LLM call traced
→Same code runs locally, on the devbox, or on your cluster
agent.py flyte-sdk

Deploy Real-time Model Endpoints and Dashboards

The same Python API serves long-running apps next to your workflows. Native integrations wrap Streamlit, FastAPI, vLLM, SGLang and Ollama — declare the app, image, resources and scaling; Flyte handles the serving plumbing.

→OpenAI-compatible LLM endpoints with model streaming and tensor parallelism
→Interactive dashboards served as long-running apps, auth included
→Scale to zero when idle — replicas and scaledown are one line of config
serve_llm.py flyteplugins-vllm

Built for the failure modes you actually hit

COMPUTE DURABILITYRecover from infra failures

OOM kills, preempted spot nodes, disappearing GPUs — Flyte is infra-aware and resubmits with corrected resources, automatically.

LOGICAL DURABILITYRecover from code failures

Retry and resume from the failed task with the exact same inputs. Never rerun the whole pipeline for one bad step.

DynamismAdapt at runtime

Branch, loop and make decisions during execution — built for agents and dynamic pipelines, not static DAGs.

ScaleAutoscale compute

Infrastructure scales up and down to match workload demand. Right-size every task; stop burning idle GPU hours.

NATIVE CONCURRENCY & PARALLELISMScale workloads

Thousands of parallel tasks and distributed jobs — map, fan-out and streaming map-reduce without rewriting your pipeline.

DevexInference, sandboxing, pro UI

Serve vLLM and SGLang apps, sandbox agent code, and debug with a UI built for production runs.

Make your existing stack durable

All integrations →
Apache SparkEphemeral clusters
RayDistributed training
PyTorch ElasticElastic multi-node
BigQueryWarehouse queries
SnowflakeSQL tasks
Weights & BiasesExperiment tracking
Plus OpenAI Agents SDK, LangGraph, CrewAI, Pydantic AI, vLLM, SGLang, MLflow, Dask, Databricks and the community plug-in registry.

Union.ai — the enterprise Flyte platform

Orchestrate, ship and scale AI systems from experiment to production: training, real-time inference, and observability on managed infrastructure.

Compare Flyte OSS and Union.ai