GPT-6 Astra in Copilot Cowork and in Foundry

When I wrote about GPT-6 Astra in Microsoft Foundry a few weeks ago, I did not yet have access. There was plenty to discuss about what the model promised, but the most interesting part was still missing: using it in my own work.

Now I can start doing that! I am writing this article with GPT-6 Astra selected in Copilot Cowork, using information I collected with Copilot as the starting material and my blog-writing skill to guide the draft. A slightly meta example: using Astra to help write about using Astra. But let’s keep the distinction clear. This is a real writing task, not yet a comparison proving that Astra is better than the other available models. The interesting question is no longer just what Astra can theoretically do. It is what changes when I put it to work.

Let’s take a closer look.

From a model announcement to a work option

For me, the biggest change is not another set of specifications. I covered those in the earlier article. It is the shorter distance between hearing about a model and using it for something I actually need to get done.

My starting point here was not an empty prompt asking for an article about GPT-6. I brought research, my previous post, a direction for the follow-up, and my own writing guidance. I also identified details that still needed checking.

That is much closer to how I want to work with AI.

Here is the context. Here is what I am trying to achieve. Help me move it forward, without pretending the missing pieces are already known.

For this article, that means separating announcements from observations, avoiding another launch recap, and leaving room for tests I have not completed yet. A polished draft that quietly invents those results would be worse than an unfinished one that shows exactly what is missing.

The writing is only part of the task. Knowing what not to include and refining the result from the AI draft is very important.

Foundry and Cowork – the model is only part of the experience

I see two different starting points here.

With Microsoft Foundry, my question is about building a solution. How should an application or agent be designed, connected, evaluated, deployed, and operated?

With Copilot Cowork, my question is about the work in front of me. What can I delegate, what context should I provide, and how much checking will the result need?

Those questions are related, but they are not interchangeable.

Selecting Astra in Cowork is not the same thing as deploying Astra in Foundry and building an application around it. The surrounding experience matters: instructions, tools, available information, permissions, and the way the work is coordinated.

That also matters when evaluating results.

If an article draft turns out well, I cannot automatically credit the model for everything. Background information provided are essential part of the setup. Likewise, a disappointing result might reveal missing context rather than a limitation in reasoning.

I want to evaluate the whole way of working, not just admire the model badge.

For my own work, Cowork is the place to start when I want help completing a substantial task in an existing work experience. Foundry becomes relevant when I need to build and operate a repeatable solution with explicit architectural and deployment choices.

This article is a starting point, not a verdict

Writing a follow-up is actually a useful first task. There is an existing argument to continue without repeating it. There are technical details with different levels of certainty. There is a personal voice to preserve. And there is a temptation to turn an enthusiastic introduction into a conclusion before the evidence exists.

For this draft, I want Astra to help with three things:

  • Find the new story. What has changed enough to justify another article?
  • Keep the boundaries visible. What is documented, what have I personally observed, and what remains untested?
  • Produce something worth editing. Not simply more text.

That last point is important. A longer response is not automatically a better result.

If I spend more time removing repetition, checking unsupported statements, and putting my own perspective back into the draft, the apparent productivity gain can disappear quite quickly.

This is also why I find the combination of a model and a reusable writing skill interesting (yes, I have a skill in Cowork that helps me write blog posts). It provides guidance about tone, structure, and how to handle uncertainty rather than explain everything again from scratch.

Astra versus Auto – make it earn the selection

For my next tests, I want to compare Auto with explicitly selected Astra.

I am not starting with the assumption that manually choosing Astra must be better. If Auto gets me an acceptable result with less effort, that is useful too. And it often is – I will usually just use Auto for most tasks.

I would start with tasks that resemble my actual work:

  • Turn research and personal notes into an article with a clear argument.
  • Compare several documents and produce a recommendation that preserves disagreements and uncertainties.
  • Build a workshop outline from a brief, with a coherent flow and practical exercises.

I want to keep the source material, instructions, and acceptance criteria consistent. I would also use separate tasks so that one attempt does not benefit from corrections made during the other.

Then I can look beyond which response feels more impressive at first glance. Did it follow the brief? Did it preserve important qualifications? Did it miss a source? How much rewriting was necessary? Was the deliverable usable?

Reasoning effort – what does this task deserve?

Cowork’s model guidance describes five reasoning-effort choices: Light, Medium, High, Extra High, and Max. Higher effort involves a trade-off in response time and Copilot Credits. Microsoft explains this under Set the reasoning effort level to balance quality, speed, and cost in Choose a model for Copilot Cowork.

That makes this more interesting than simply choosing a model.

My proposed comparison is to run the same substantial task at Medium, High, and Max, then record:

  • Time to an initial deliverable.
  • Credits consumed, where usage can be reliably attributed to the task.
  • Clarification questions and correction rounds.
  • Factual errors or missed requirements.
  • My own review and editing time.
  • Whether I accepted the result, revised it, or started again.

If I cannot measure task-level credit consumption reliably, I will say so rather than estimate it.

The question I want answered is simple: does the extra reasoning (and cost) reduce the total effort needed to get an acceptable result?

A slower run could be worthwhile if it saves substantial review. A faster one could be the better choice when the task is straightforward. And Max should have to demonstrate its value—not become my default because it sounds reassuring.

The same applies to cost. A cheaper individual run is not necessarily a cheaper completed task if it needs several retries and extensive manual correction. On the other hand, spending more does not automatically buy an improvement that matters.

I want to measure the whole journey to something I can use.

A note for us in Europe – documented, but not yet visible in my Foundry EUR Datazone

In my first Astra post, I noted that the published deployment options were Global and U.S. Data Zone. There is now an update for us in Europe.

Microsoft’s September 22 article, GPT-6 Astra, Sol and Luna: For production agents in Microsoft Foundry, explicitly includes Astra in EU Data Zone availability.

The Foundry Models region availability page also listed Astra for the EU Data Zone when I checked.

Of course, I went to look in my own Foundry. And… not yet at least for me!

On the morning of September 23, 2026, Astra was not available for me to deploy to the EU Data Zone. I could deploy it to Global Standard or U.S. Data Zone.

GPT-6 Sol, on the other hand, was already available to deploy to the EU Data Zone.

It looks to me like EUR Datazone availability for Astra may still be rolling out. Given the published availability, I expect this to happen soon.

If processing location matters for your project and customer, check the model and deployment type in your own subscription before committing to an architecture or promising availability to a customer.

One more distinction: Foundry deployment options do not establish where a Copilot Cowork task using GPT-6 Astra is processed. Those arrangements need to be checked separately. At this moment GPT-6 Astra in Cowork is a preview model and runs using OpenAI subprocessor, which doesn’t define the datazone. Now that Astra is in Foundry, this hopefully changes soon.

What did writing this article cost?

I have also been following the Copilot Credits consumed while working on this article, using /cost in Cowork.

Before adding this cost information and further updates, the total was 232 Copilot Credits. That was much less than I expected.

This is an observation from this particular writing session with GPT-6 Astra selected—not a fixed price for an article, and not a comparison with Auto or another model. It also excludes the updates made after that reading.

Still, it gives me a concrete starting point. Instead of only discussing what deeper reasoning might cost, I now have an actual usage figure from my own work.

The next question is how that consumption compares with the quality of the result and the time I spend reviewing and editing it. Credits are one part of the cost. My own time is another.

From a good answer to accepted work

I am excited that I can now include Astra in my own Cowork experiments. But I do not want the conclusion of this follow-up to be that a new model appeared and therefore everything improved. The useful conclusion will come from the work.

For this article, success means a follow-up that adds something new, keeps uncertain details visible, and takes less effort to finish without lowering the quality. For another task, the acceptance criteria will be different.

That is where I want to focus next: the cost and reliability of a completed, reviewed, accepted outcome.

Not just what the model produced but what I could actually use.

That is a more useful conversation about the future of work than choosing a favourite model and assuming it belongs in every task. We need to learn where deeper reasoning helps, where a lighter approach is enough, and where human judgement remains essential.

Now, on to the testing!

Leave a comment

This site uses Akismet to reduce spam. Learn how your comment data is processed.