Dynatrace signed a definitive agreement on August 13 to acquire Arize, an AI observability company, for $915 million in a cash and stock transaction. The deal is designed to expand Dynatrace's reach into the AI developer community and secure a position earlier in the AI application development lifecycle. The purchase is a bet on lifecycle position rather than feature parity, and it comes as every major observability incumbent already ships AI evaluation tools.
The terms are roughly $815 million in cash plus replacement equity awards for Arize employees joining Dynatrace. Dynatrace plans to fund the acquisition from cash on hand or its existing credit facility. Closing is expected this quarter or early next, subject to regulatory review. Dynatrace did not disclose Arize's revenue, and the multiple on that figure is impossible to calculate from public information.
Why Dynatrace Is Paying for Position
Dynatrace was already shipping evaluation before this deal. Its existing AI Observability app traces gen_ai spans, scores live production responses with LLM-as-a-judge evaluators, and detects drift in scores over time. In June, Dynatrace open-sourced dt-evals, a command-line tool that pulls recent gen_ai spans and scores them with an LLM judge. The results are written back as business events linked to the source trace. Documentation for dt-evals lists more than 10 built-in judge evaluators plus statistical drift detection against a rolling baseline of earlier scores.
What Dynatrace lacked was a relationship with the AI engineers who choose an evaluation harness. Tooling choices get made while an application is still being written, months before anything reaches an operations team. By the time an application reaches production, the instrumentation library, the trace schema, and the evaluator definitions have already been chosen. Dynatrace is paying to be in the room when those decisions are made.
Arize is strongest in the half of the lifecycle that runs before an application has production traffic to score: experiments, datasets, prompt iteration, and pre-release evaluation. The company reaches developers through Phoenix, a self-hostable tracing and evaluation project, and enterprises through the commercial AX platform. Dynatrace's rationale that the deal expands its reach into the developer community is the part that holds up. The decisive difference between Dynatrace and competitors is where the tooling decision starts.
The Competitive Landscape Is Already Crowded
Datadog, Splunk, and New Relic already have AI evaluation features. Datadog traces LLM and agent applications, tracks token usage and cost, and supports managed and custom LLM-as-a-judge evaluations attached to individual spans. Splunk's AI Agent Monitoring runs platform-side and instrumentation-side evaluations covering hallucination, bias, relevance, sentiment, and toxicity. Splunk's documentation says an agent gets flagged when fewer than 80% of evaluations pass for a metric. New Relic has its own AI monitoring across models, traces, cost, and performance.
Those features emerged as extensions of operations and platform engineering relationships. Datadog, Splunk, and Dynatrace sell into operations and platform engineering. Arize built from the opposite end, with Phoenix as a free local project that AI engineers adopt long before a procurement conversation exists. In a category where every incumbent already has the features, position is the right thing to buy.
A single agent run may involve a model request, a document retrieval, several tool calls, and a final response, each landing as a span within a single trace. An evaluator can be deterministic code, a human annotation, or another model acting as a judge. An evaluation output is a score against a particular rubric and evaluator, not a verdict on truth. Whether an AI system ran and whether it produced an acceptable result are now two separate operational questions.
Phoenix, OpenInference, and the OpenTelemetry Question
Phoenix uses OpenInference as its native semantic format rather than OpenTelemetry conventions for generative AI. Traces arriving from other libraries get translated into OpenInference so they display consistently in Phoenix. Arize AX now normalizes compatible gen_ai attributes into OpenInference fields during ingestion, removing the need for a client-side conversion processor. Arize treats both OpenInference and OpenTelemetry conventions as first-class, and it expects them to converge as the OpenTelemetry spec stabilizes.
That architectural choice matters for interoperability. OpenTelemetry is the broader observability standard, and generative AI conventions are still settling. Arize's bet is that OpenInference and OpenTelemetry will converge, and the company has positioned its tooling to work with both. For Dynatrace, that means the acquisition brings a semantic layer that can translate between the two worlds.
The licensing question is more delicate. Arize describes Phoenix as open source, and the main repository ships under the Elastic License 2.0. That license permits broad use and self-hosting, but it restricts anyone from offering the software itself as a hosted or managed service. The Elastic License 2.0 is not approved by the Open Source Initiative. The distinction matters to exactly the developers whose trust Dynatrace is paying for. A developer who expects OSI-approved open source may find the restriction surprising.
Stay ahead of the AI curve
The most important updates, news, and content — delivered weekly.
No spam. Unsubscribe anytime.
The Financial Shape of the Deal
Dynatrace reported $2.14 billion in ARR and a 29% non-GAAP operating margin for the June quarter. The acquisition is expected to add roughly 200 basis points of accretion to ARR growth in the coming fiscal year. Dynatrace also guided to a 175 basis-point dilution in the non-GAAP operating margin, with expansion expected the year after. The company is accepting a year of margin dilution and spending close to a billion dollars while the category is still forming.
Reverse-engineering an Arize ARR figure from the accretion guidance does not work. The guidance describes an effect on Dynatrace's own growth rate, including the timing of the deal. The company did not break out Arize's revenue, and the public information is insufficient to estimate a multiple.
Jason Lopatecki, co-founder of Arize, will join Dynatrace at closing and continue to lead the Arize team, reporting to Dynatrace CEO Rick McConnell. Aparna Dhinakaran, the other co-founder, will also join Dynatrace at closing. The deal keeps the founding team in place, which matters for continuity with the developer community Arize has built.
The Risks of Probabilistic Evaluation
The evaluations themselves are probabilistic. An evaluator can be deterministic code, a human annotation, or another model acting as a judge. When the judge is a model, the output is a score against a particular rubric, not a verdict on truth. An evaluator can disagree with a human reviewer, and it can drift when its underlying model version changes. A monitoring system with an LLM judge has a second model embedded in its control loop.
Running an LLM judge still costs model tokens paid to the provider. Arize AX meters span volume and ingested data instead of evaluations, and it lists evaluations, experiments, and human annotations as unlimited across its Free, Pro, and Enterprise plans. Tracing the evaluator's own execution consumes the same span allowance. An evaluation score behaves like a sampled quality indicator rather than an HTTP status code.
Running evaluations in production calls for versioned evaluators, held-out test sets, periodic human calibration, and an explicit threshold before a score blocks a deployment. Without those controls, a drifting judge can quietly change what passes and what fails. The operational discipline required to run probabilistic evaluation at scale is not something a monitoring platform provides by default.
What Buyers Should Ask
Merging evaluation into an observability platform does not resolve ownership boundaries, though it does put them on one screen for the first time. The first question for a buyer is ownership of instrumentation. Whoever controls the instrumentation library controls the trace schema, and that decision gets made before an operations team is involved. The second question is about evaluation economics. Who pays for the model tokens, and who owns the span allowance when the evaluator runs? The third question is ownership of the quality signal. If an evaluation score blocks a deployment, who has the authority to override it?
Dynatrace is taking a calculated risk on lifecycle position rather than on features. The purchase buys a seat at the table where AI engineers make their earliest tooling decisions. Whether that position translates into durable revenue depends on whether the developer community accepts the licensing terms and whether the evaluation tools hold up under production load. The company is betting that being present at the start of the lifecycle is worth more than matching a feature list.
The deal is expected to close this quarter or early next, subject to regulatory review. Dynatrace plans to fund the acquisition from cash on hand or its existing credit facility. The company guided to roughly 200 basis points of accretion to ARR growth in the coming fiscal year and a 175 basis-point dilution in the non-GAAP operating margin, with expansion expected the year after. Dynatrace reported $2.14 billion in ARR and a 29% non-GAAP operating margin for the June quarter. The $915 million price tag, with $815 million in cash, is a bet that the AI evaluation category is still forming and that the developer relationship is the asset worth buying.

