Skip to main content
Hyperagent
ConceptsAgents

Models and compute

Choose from a wide model catalog and match reasoning, speed, and cost to the size of the job.

Hyperagent supports a wide range of models, from fast, efficient models for routine work to frontier models built for deep reasoning. That range lets you choose the right level of thinking for the size and stakes of the job instead of giving every agent the same model.

Match the model to the size of the job

A one-line triage check does not need the same reasoning budget as a dense research brief. Give the recurring role a sensible default, then override it when one conversation needs more depth or more speed.

Match the model to the job

Model choice changes the quality, latency, and cost of every run. Start with the work the agent owns, then choose enough capability to do that work reliably.

Lightweight

Routine, frequent work

Classification, formatting, extraction, and simple recurring checks reward speed and a lower cost per run.

Balanced

Everyday agent work

Research, drafting, analysis, and multi-step tool use need dependable judgment without paying for the strongest model on every turn.

Deep reasoning

Large or consequential jobs

Ambiguous research, difficult strategy, complex debugging, and external final work benefit from a frontier model and more reasoning effort.

Where you choose a model

The same model choice appears in four places. Choose the surface that matches the scope you want to change.

Named agent

Model and limits

Sets the agent's default model, runtime, subagent model, turn timeout, and auto-compact threshold for new work.

Open thread

Header model chip

Changes the model for the current conversation without changing the named agent's default.

Thread settings

Model & compute

Holds the same thread override beside Fast inference and Reasoning effort, so the model and its compute controls stay together.

No named agent

Agent defaults

Sets the model a home thread starts on when no agent is bound. Open Agent defaults to change it.

Set the usual model, change it when needed

The model saved on an agent is the starting point for its work. You can change the model for one conversation without changing what the agent uses everywhere else.

Agent setting

The usual model for new threads

Choose the model that fits the role's recurring work. Every new thread on that agent starts with this model unless you change it inside the thread.

Thread setting

A different model for this conversation

Change the model from the thread header or Model & compute. That thread keeps the new choice; the agent's usual model and its other threads stay unchanged.

A home thread that is not connected to a named agent starts with your Agent defaults. Once any thread begins using a model, changing the agent's usual model does not rewrite that existing conversation.

Latest aliases and pinned versions

The picker opens in two sections, and which one you choose decides whether the agent moves with upgrades or stays put.

The model picker open, with a Latest models section holding maintained Latest aliases above an All models section of pinned versions grouped by provider.
The picker opens in two sections: maintained Latest aliases on top, pinned versions grouped by provider below, with Latest and Experimental badges.

Latest models

Default

Maintained aliases such as Latest (Opus) or Latest (GPT) that track the newest version in a family. Pick one and the agent receives the family's normal upgrades with no configuration change. Hover the section for the note: We'll automatically use the latest model in the selected family when you create a new thread.

All models

Pinned numbered versions, grouped by provider. Pick a pin when you need repeatability, such as an evaluation or a controlled comparison, and the model stays exactly where you set it until you change it.

When an alias is selected, the model chip carries a Latest badge. Experimental models carry an Experimental badge, with the tooltip Early access, behavior and availability may change.

Supported model catalog

Hyperagent supports models from several labs, so you can match the model to the agent's job instead of rebuilding the agent around one provider. The catalog changes as models are validated; the picker is the source of truth for what is available to your account.

Anthropic

ModelBest fitContext
Fable 5.1Sustained multi-turn workflows, long planning, and 4x cheaper cache reads1M
Fable 5Highest-capability work when quality matters more than cost1M
Opus 5Complex agents, difficult reasoning, and careful final output1M
Opus 4.8Earlier flagship for complex work; supports Fast inference1M
Opus 4.7Earlier Opus model for complex work1M
Opus 4.6Earlier Opus model for complex work1M
Sonnet 5Balanced everyday agents at lower cost than Opus1M
Sonnet 4.6Earlier balanced Sonnet model for everyday work1M
Haiku 4.5Quick, simple tasks and high-volume lightweight work200K

OpenAI and Google

ModelBest fitContext
GPT-6 AstraOpenAI's frontier reasoning model for complex agentic analysis950K
GPT 5.6 SolDeep reasoning model with image support950K
GPT 5.6 TerraBalanced OpenAI option for everyday agents950K
GPT 5.6 LunaFast, low-cost OpenAI option950K
GPT 5.5Earlier OpenAI flagship950K
Gemini 3.8 FlashHigh-throughput, fast Google model with 1M context1M
Gemini 3.6 FlashFast structured execution, monitoring, research loops, and coding1M
Gemini 3.5 FlashEarlier fast Google model for everyday work1M

Open and specialist models

ModelBest fitContext
GLM 5.3 FlashFast multimodal open model for coding and agentic work900K
GLM 5.2Low-cost everyday agent work1M
GLM 5.2 FastFaster GLM output for high-volume work1M
Kimi K3Moonshot's highest-capability open model900K
Kimi K3 FastFaster Kimi K3 output when latency matters900K
Kimi K2.6Low-cost everyday work on an open model230K
Qwen 3.8 MaxFrontier coding, research, and long-horizon work1M
Qwen 3.7 PlusCost-aware general work and coding230K
DeepSeek V4 FlashHigh-volume coding and tool work at the lowest cost1M
DeepSeek V4 ProDeep reasoning on an open model1M
MiniMax M3Low-cost search and synthesis across a large internal knowledge base1M
Fugu UltraCoordinated expert-agent work on complex tasks1M
Grok 4.6Longer research and analysis that ends in a decision-ready brief500K
Grok 4.5Conversational agent work500K
Muse Spark 1.1Creative agent work1M
InklingCompact, low-cost open-model work230K

If a model you need is missing, share both the model and the job you want it to perform with the Hyperagent team. The use case helps the team evaluate quality, runtime compatibility, and demand.

Model & compute controls

Inside a thread, Model & compute gathers the model and the two dials that shape how it runs. Reasoning effort decides how hard the model thinks; Fast inference trades cost for speed on the models that support it.

The Model and compute section of a thread, showing the model row with a Latest badge, the Fast inference toggle, Reasoning effort set to Medium, and the subagent model row
Model & compute in a thread's settings: the model with its Latest marker, Fast inference, Reasoning effort, and the subagent model.

Reasoning effort

Reasoning effort sets how much thinking a model does before it answers. More effort buys deeper reasoning on hard problems; less effort returns faster and costs less. The default is Medium, shown as Balanced.

LevelReads as
LowFast responses
MediumBalanced
HighDeep reasoning
Extra highDeeper reasoning
MaxMaximum capacity

Reasoning controls depend on the model

Each model shows only the effort levels it supports, so its picker may contain a subset of these five. When a model has no configurable reasoning, the effort picker does not appear.

Fast inference

Fast inference runs an eligible model at higher throughput. Its helper text states the tradeoff plainly: Faster output, billed at 2x token cost. Reach for it on interactive replies and tight loops where latency matters, and leave it off for deep analysis and final writing where standard speed is the better fit.

The toggle appears only on the models that expose it, currently Opus 4.8 and Opus 5. If you don't see the toggle on the model you're using, that model doesn't offer Fast inference in Hyperagent.

Subagent model

Subagents are short-lived workers a run spins up to handle parallel legwork on its own. Subagent model sets the default model those workers use.

Sonnet

Default

Balanced capability and cost. The default subagent model, shown as Default (Sonnet 4.6).

Opus

Frontier reasoning for delegated work that needs the strongest model.

Haiku

Fastest and lowest cost, for high-volume lightweight legwork.

Parent model

Match the parent agent's own model instead of a fixed tier.

Runtime

The Runtime control on the agent's Model and limits card decides which execution runtime serves the model. It defaults to Default chooses the runtime from the selected model, which is the right setting unless you have a specific reason to change it.

Turn timeout

Turn timeout caps how long one turn may keep running. It applies to the current turn, not the lifetime of the thread: later turns can continue in the same conversation.

Use a longer timeout for deep research, browser work, coding, or other jobs that may need sustained tool use. A shorter timeout gives routine or unattended work a firmer stop. If the limit is reached, Hyperagent ends that turn rather than letting it run indefinitely, and the thread keeps the record of where it stopped.

Auto-compact threshold

Every model has a finite context window. Auto-compact threshold decides how full a thread's active context may become before Hyperagent compacts it automatically. Compaction summarizes older conversation so the agent has room to continue; it does not remove the transcript from the thread.

A lower threshold compacts sooner and leaves more headroom for upcoming work. A higher threshold keeps more of the raw conversation active before compacting. For a one-time compaction, open the model menu in the thread and choose Compact context now, or enter /compact.

The thread's Thread Context Document survives compaction and keeps the facts, corrections, decisions, and plan tasks the agent recorded for the current job. See how a thread keeps working context.

Live Mode model

Live Mode runs each scheduled check on its own heartbeat model, chosen separately from the thread and defaulting to the latest Sonnet. This lets a frequent check use a faster, less expensive model without changing the model used for the agent's other work.

When a model becomes unavailable

If a model is removed from your account, it leaves the picker and any new run that resolves to it errors with not available for your account. Latest aliases migrate gracefully: when a family is retired behind an alias, stored selections resolve to the model now serving that family, and the thread shows the model actually in use.

FAQs