Adventures with Copilot Studio: Choosing the Right Model and Harness

I wanted to put together a cheat sheet for the models available in Copilot Studio. Not just a list of names, but something that helps answer a more useful question: which one would I use for the work I need the agent to do?

There are really two decisions here. Which harness fits the job? And which model fits the reasoning?

Picking the newest model does not answer both questions.

📌 Before we get started

Checked against Microsoft Learn on October 8, 2026. This verifies the published documentation—not which models are enabled in your tenant. The agent tables use United States commercial-cloud availability. The prompt table now includes United States release status. Region, environment, rollout, and administrative settings still matter.

🧭 First, which harness do I need?

The model gets most of the attention. But the harness matters just as much. It is the runtime around the model, and it affects how the agent works with tools, files, and the process you want it to complete.

🧭 Harness💡 Choose it when…📖 Business scenario🤖 Model-selection reference
Standard harnessYou want authored topics, controlled branching, and repeatable processes.A plant employee requests a replacement device; the agent collects required fields and follows an approved routing path.Standard harness models and prompt-builder models.
GitHub Copilot harnessWork requires adaptive planning across tools, systems, and documents.An accounts-payable agent investigates an invoice mismatch, gathers supporting records, prepares an exception package, and requests approval.GitHub Copilot harness models.
Copilot chat harnessThe main goal is employee access to enterprise knowledge inside Microsoft Copilot Chat.A new hire asks where to find onboarding requirements and relevant SharePoint guidance.The harness overview says it uses current chat models; it does not provide a selectable, versioned model inventory. Do not assume the other harnesses’ pickers apply.

Scenarios and selection recommendations are illustrative guidance, not deployed customer examples.

One thing worth calling out: The GitHub Copilot harness in Copilot Studio is not the GitHub Copilot service. Similar name, different context.

⚙️ Which models can I use with the standard harness?

For a predictable service conversation, I would start by evaluating the default model against the requests the agent actually needs to handle. Then compare alternatives using the same test cases.

Selection surface: Agent Overview → Model. These choices concern orchestration; they are not automatically the same choices as prompt builder.

🤖 Model🏷️ Category🌎 US availability💡 Suggested business use / action
GPT-5.5 ChatGeneralDefaultStarting candidate for grounded employee help, customer-service conversations, and routine tool use.
GPT-5 ChatGeneralGACompare against the default using an existing agent’s regression test set.
GPT-4.1GeneralGARetain or evaluate where existing prompts and workflows have already been validated.
Claude Sonnet 4.6GeneralGAAlternative candidate for grounded service conversations and drafting; validate against the same cases.
Claude Opus 4.6DeepGACandidate for complex policy interpretation or multi-document exception analysis.
Claude Opus 4.7DeepGAAnother deep-model candidate for difficult cases; measure rather than assume superiority.
GPT-5 ReasoningDeepPreviewNonproduction evaluation of nuanced troubleshooting or policy analysis.
GPT-5 AutoAutoPreviewNonproduction evaluation of a help desk with both simple and complex requests.
GPT-5.3 ChatGeneralExperimental; early-access environmentControlled evaluation only.
GPT-5.4 ReasoningDeepExperimental; early-access environmentControlled evaluation of complex reasoning only.
GPT-5.5 ReasoningDeepExperimental; early-access environmentControlled evaluation of complex reasoning only.
Grok 4.1 Fast (Non-reasoning)GeneralExperimental; early-access environmentEvaluation only; review Microsoft’s additional safety warning before testing.
Mistral Medium 3.5GeneralExperimental; cross-geoEvaluation only; review data movement and provider requirements.
GPT-4oGeneralRetired in commercial regionsNot a new commercial deployment choice. See government note below.
Claude Sonnet 4.5GeneralRetiredPlan migration, not a new deployment.

Suggested uses are category-based starting points, not model-specific benchmark claims.

Government exception: The same reference lists GPT-4o as the default for GCC, GCC High, and DoD. Do not carry the commercial-cloud inventory into government environments.

🤖 What about the GitHub Copilot harness?

If the job involves investigating a problem, coordinating several tools, or producing files, this is the harness I would evaluate. That still does not mean every request needs a Deep model.

🤖 Model🏷️ Category🌎 US availability💡 Suggested business use / action
GPT-5 ChatGeneralGACandidate for straightforward document drafting and routine tool-supported work.
GPT-5.5 ChatGeneralGAGeneral-workload candidate; validate quality and consumption on representative tasks.
GPT-6 AstraDeepGACandidate for cross-system investigation, exception handling, and evidence synthesis.
Claude Sonnet 4.6GeneralGAAlternative candidate for drafting and routine knowledge-plus-action tasks.
Claude Sonnet 5GeneralGA; early-access environmentGeneral-workload candidate where the early-access environment exposes it.
Claude Fable 5DeepGA; early-access environmentComplex-workflow candidate where available.
Claude Fable 5.1DeepGA; early-access environmentComplex-workflow candidate where available; compare on your own evaluation set.
Claude Opus 4.8DeepGACandidate for document-heavy analysis and multi-step investigations.
Claude Opus 5DeepGAAnother deep-workload candidate; select on measured business outcomes.
GPT-5.6 ReasoningDeepExperimental; early-access environmentControlled evaluation only.
Mistral Medium 3.5GeneralExperimental; cross-geoControlled evaluation only.

The availability table does not label a default model. Early-access availability is an environment condition, not synonymous with experimental status.

A quick regional gotcha: The reference shows no availability for Sonnet 5, Fable 5, or Fable 5.1 in Australia or Saudi Arabia. Many other entries require cross-geo processing outside the US. Check the region matrix before recommending a model.

✍️ Wait, does prompt builder use the same model?

This is where it is easy to mix things up. The primary agent model and the model selected for a prompt are separate choices. A prompt might just extract an asset ID or classify a ticket. It does not necessarily need the same model doing the orchestration.

🤖 Prompt model💳 Published rate tier🌎 US release status💡 Suggested task
GPT-4.1 miniBasic; defaultGAExtract purchase-order references, classify tickets, or summarize a short handoff.
GPT-4.1StandardGAProduce structured summaries from more demanding source material.
GPT-5 chatStandardGADraft a customer response from approved facts.
GPT-5.3 chatStandardExperimental — evaluation onlyEvaluation only: compare general-purpose prompt outputs. Not a production recommendation.
GPT-5 reasoningPremiumGAAnalyze an invoice exception with multiple constraints.
GPT-5.2 reasoningPremiumExperimental — evaluation onlyEvaluation only: test complex comparisons and reasoning. Not a production recommendation.
Claude Sonnet 4.6StandardExperimental — evaluation onlyEvaluation only: test general-purpose prompt tasks. Verify this surface’s status before production use.
Claude Opus 4.6PremiumExperimental — evaluation onlyEvaluation only: test reasoning prompt tasks. Verify this surface’s status before production use.
Grok 4.1 Fast (Non-reasoning)StandardExperimental — evaluation onlyEvaluation only; additional safety review required. Not a production recommendation.

A billing tier does not establish production readiness. Rate tiers are not dollar quotes or guarantees of cost per business transaction.

⚠️ Here is a gotcha: Check release status for the specific selection surface—not just the model name. Claude Sonnet 4.6 and Claude Opus 4.6 are experimental in prompt builder, even though agent model tables list them as GA. Treat these prompt choices as evaluation only. The tables above should not be read as a shared release-status list across surfaces.

The primary-agent documentation also identifies separate deep-reasoning and generative-response settings. Their inventories are not established by the tables above.

🏭 What does this look like in a real business scenario?

A list of model names is helpful. But what would I actually do with them? Here are a few places I would start. These are examples to evaluate, not a claim that one model wins every time.

💼 Business request🧭 Harness🤖 Starting model approach🛡️ Why / controls
“Where is the approved supplier onboarding policy?”Copilot chatCurrent chat models supplied by that experienceKnowledge access, not a long-running transaction. Keep source permissions and citations intact.
“Submit a standard IT equipment request.”StandardGPT-5.5 Chat plus authored stepsA known process benefits from required fields and controlled routing.
“Explain my order status and open a case if it is overdue.”StandardGPT-5.5 Chat; evaluate other GA General modelsGround answers in actual order records; define the condition for creating a case.
“Classify these maintenance tickets and extract asset IDs.”Standard process with a promptGPT-4.1 mini as the initial prompt candidateA bounded transformation can be tested against an expected output schema. Use rules for identifier validation.
“Investigate this invoice/PO mismatch and prepare an approval package.”GitHub CopilotEvaluate GPT-6 Astra, Claude Opus 4.8, or Claude Opus 5Multi-system evidence and documents; require human approval before payment or ledger changes.
“Compare supplier agreements and flag conflicting obligations.”GitHub CopilotEvaluate GA Deep candidatesDocument-heavy synthesis; require clause citations and legal review rather than autonomous legal decisions.
“Build a weekly manufacturing operations briefing from ERP records and approved reports.”GitHub CopilotGeneral for straightforward assembly; evaluate Deep for exception analysisFile creation and synthesis; validate every metric against its source. Schedule only after testing.
“Diagnose a recurring quality issue across inspection, maintenance, and supplier records.”GitHub CopilotEvaluate GA Deep candidatesSeveral evidence sources and hypotheses; label uncertainty and keep safety-critical actions with authorized personnel.
“Test adaptive handling of mixed-complexity help-desk questions.”Standard, nonproductionGPT-5 Auto, PreviewAn experiment—not the production default. Keep a validated GA alternative.

💡 So, where would I start?

  • A question about company knowledge: start with the knowledge experience, not an autonomous workflow.
  • A fixed, repeatable business path: start with the standard harness and explicit process controls.
  • A goal requiring investigation, tool coordination, or a document deliverable: evaluate the GitHub Copilot harness.
  • A small classification or extraction step: start with a mini prompt candidate and validate its output.
  • Several interacting constraints: evaluate GA Deep models against representative cases.
  • Mixed workloads: test whether one GA General model is sufficient before adding complexity.
  • Production: select a GA model that meets region, policy, quality, latency, and consumption requirements. Do not select Preview or Experimental models for production.

These are implementation recommendations, not model-specific performance guarantees.

✅ What should I check before putting this into production?

  • Confirm harness, selection surface, cloud, environment region, and release cycle.
  • Check the actual model picker; documentation availability does not establish tenant enablement.
  • For external models, check both Power Platform environment/environment-group controls and provider access in Microsoft 365 admin center.
  • Review cross-region data movement and preview/experimental controls separately.
  • Test representative successful cases, missing data, ambiguous requests, denied permissions, and failed tools.
  • Measure task completion, factual grounding, response time, and credits consumed; do not choose solely by model name.
  • Keep human approval for consequential financial, legal, production, or safety decisions.
  • Re-test when models, defaults, prompts, tools, or policies change.

My take: Start with the business problem, not the newest model name. Use the simplest harness and GA model that can do the job reliably. If deeper reasoning improves the outcome in your tests, great. If it does not, keep it simple.

📚 References

Harnesses in Copilot Studio — “GitHub Copilot harness,” “Standard harness,” “Copilot chat harness,” and “Compare harnesses.” Updated October 1, 2026. https://learn.microsoft.com/en-us/microsoft-copilot-studio/harnesses-overview

Select a primary AI model for your agent — “Standard harness availability,” “US Government availability,” model categories, selection, and administrative controls. Updated September 18, 2026. https://learn.microsoft.com/en-us/microsoft-copilot-studio/authoring-select-agent-model

Model availability and controls — GitHub Copilot harness inventory, regional availability, release types, and administrative controls. Updated October 2, 2026. https://learn.microsoft.com/en-us/microsoft-copilot-studio/agents-experience/authoring-agent-model-availability

Change the model version and settings — Prompt builder inventory, rate tiers, and release-stage caveats. Updated August 4, 2026. https://learn.microsoft.com/en-us/microsoft-copilot-studio/prompt-model-settings

Prompt model availability by region and updates — “Public availability,” United States release status. Updated August 3, 2026. https://learn.microsoft.com/en-us/microsoft-copilot-studio/prompt-model-availability

Availability changes frequently. Recheck these references and your environment before sharing this as deployment guidance.

Written with the help of AI.


Discover more from Matt Ruma

Subscribe to get the latest posts sent to your email.

Leave a Reply

Your email address will not be published. Required fields are marked *