Product
Accelerating scientific research with GPT-6 Astra
Astra brings stronger scientific reasoning and computer use to the systems where biopharma research happens.
by Enjamb Team

GPT-6 Astra is now available in Enjamb. Enjamb puts Astra's scientific reasoning and computer use to work across the research process, from investigating a question to preparing the evidence for the next experiment.
In our internal testing, we put Astra through the workflows described in this post: scientific evidence reviews, data analysis, and computer use across connected systems. These are the assignments that bring the model's capabilities into the day-to-day work of a research team.
That path contains an enormous amount of work. A scientist has to find the relevant literature, reconcile it with internal results, prepare data, run analyses, and decide whether a finding holds up. Each step can involve a different system. Astra gives Enjamb a more capable model to carry that work forward, with access to the company's tools and scientific context.
A substantial jump on scientific workflows
Terminal-Bench Science is particularly relevant to our work. Built with practicing researchers, its first release contains 70 scientific tasks and checks the artifacts agents produce, including analyses, simulations, and code. It tests whether an agent can carry out demanding research work.
OpenAI reports a 64.6% score for Astra, up from 22.4% for GPT-5.6 Sol: a 42.2 percentage-point gain. The release also shows stronger mathematical reasoning and computer use.
| Benchmark | GPT-5.6Sol | GPT-6Astra | Gainpoints |
|---|---|---|---|
| Terminal-Bench Science 0.1Scientific workflows | 22.4% | 64.6% | +42.2 |
| FrontierMath Tier 4 (v2)Advanced mathematics | 83.0% | 97.6% | +14.6 |
| OSWorld 2.0v2026.08.08 · offline set · partial score | 65.7% | 72.6% | +6.9 |
| ScreenSpot-ProNo tools | 76.9% | 92.7% | +15.8 |
OpenAI-reported scores at maximum evaluated effort. Evaluation environments and tools can differ from production. Gains are percentage points, not measures of research speed or Enjamb connector performance.
Source: OpenAI's Astra evaluationsThese gains are not uniform across science: LifeSciBench moves from 59.9% to 60.3%. The strongest evidence here is for executing scientific workflows, rather than a similar leap on every biology task.
Shortening the loop from evidence to experiment
Consider a discovery team deciding which compounds to take into the next assay. The decision depends on published evidence, prior experiments, assay conditions, and fresh results. Before anyone can compare candidates, that material has to be assembled and made comparable.
In Enjamb, a scientist can give Astra a connected assignment: compare the new results with the team's existing evidence, examine differences in methods and endpoints, run the specified analysis, and prepare a recommendation for review. The output can include an evidence table, the analysis code and figures, and the assumptions behind the proposed next experiment.
Keeping the evidence and analysis together lets a scientist move directly to the next question. Ask for a sensitivity analysis, examine a conflicting paper, or test an alternative explanation using the inputs already assembled. Each investigation builds on the work that came before it, with the methods and supporting evidence available for review.
The scientist still decides what the evidence supports and which experiment to run. Astra's role is to make the preparation and analysis around that decision more capable, so a team can explore more possibilities between experimental cycles.
Better computer use makes connected workflows more powerful
Enjamb's connectors give agents access to scientific systems, enterprise applications, and the internal tools a company relies on. Astra strengthens what the agent can do with that access. Better reasoning helps it interpret the information a connector returns; better computer use helps it work through supported interfaces when an assignment also involves interacting with software.
That distinction matters in a research environment. An API may expose the underlying records, while a specialized analysis tool or internal portal provides the screen a scientist actually uses. Where computer interaction is enabled, the agent needs to find the right controls, follow a sequence of actions, and check the resulting state.
ScreenSpot-Pro tests locating interface elements in professional software. OSWorld 2.0 tests longer computer workflows across applications. Together, they probe two practical requirements: selecting the right thing on screen and keeping a larger task on track.
In a connected Enjamb workflow, the assignment extends from retrieving evidence to working with it: opening the relevant record, applying a requested filter, exporting the needed result, and carrying it into an analysis or document. Astra brings stronger reasoning and computer-use capabilities to these steps across the team's enabled tools. Enjamb supplies the connections, company context, and permissions that make those capabilities useful inside the research workflow.
Scientific context makes the capability useful
A useful research agent needs to understand how a team works: which assay definitions it uses, which protocol version applies, and why a previous candidate was deprioritized. Enjamb's Company Brain brings that permitted context into the assignment, alongside the evidence available through its connectors.
The resulting work stays inspectable. Sources, analysis outputs, and open questions remain associated with the run. Access follows the requester, and actions that require approval go to the responsible reviewer. A team can move an investigation forward while keeping the evidence behind its decisions visible.
More room for the next scientific question
Enjamb brings scientific evidence, analysis, and company tools into one working context. With Astra available, teams can apply its capabilities to the work that takes a promising question through to an analysis ready for scientific review.
Start with a real workflow: a candidate comparison, an evidence review, or a recurring analysis that crosses several systems. We can walk through how Astra fits into that work and what your team should measure to judge the result.


