Skip to main content

Blog

Page 28 of 61

All Articles

Insights on AI, machine learning, and technology strategy

Two glass observability lanes show infrastructure signal fibers and quality sample cubes for a production AI workflow.
AI Development·

Your AI dashboard needs a quality lane, not just GPU charts

A green inference dashboard can still miss the failure that matters: the model is fast, available, and wrong. Production AI teams need to monitor both infrastructure quantity and output quality.

7 min read
A glowing glass bottle diorama with paper-like artifacts, crystals, and branching light paths on a dark lab table.
AI Development·

Codex is moving AI coding agents into the customer feedback loop

Braintrust and Endava show a more useful pattern for AI coding agents: faster movement from customer request to preview branch, working spec, sandbox run, or reviewable delivery artifact.

8 min read
Glass evidence blocks, trace ribbons, and a mechanical test fixture representing AI agent evaluation receipts.
AI Development·

OpenAI's eval playbook makes the harness part of the result

A 92% success rate is not enough to approve an AI agent pilot. Teams need to know what tools, retries, prompts, budgets, safeguards, and receipts produced the score.

8 min read
The weekly workflow audit: how to find the first safe AI pilot
Small Business AI·

The weekly workflow audit: how to find the first safe AI pilot

A practical weekly workflow audit helps small-business teams find the first AI pilot that is repeated, reviewable, reversible, and safe enough to learn from.

8 min read
Your AI Agent Needs a Bug Cemetery, Not Another Demo
AI Development·

Your AI Agent Needs a Bug Cemetery, Not Another Demo

AWS Bedrock AgentCore datasets point to a practical habit for reliable agents: turn production failures into versioned regression tests with locked inputs, expected tool calls, assertions, and CI gates.

7 min read
AML alert triage shows the real shape of enterprise AI automation
Industry Insights·

AML alert triage shows the real shape of enterprise AI automation

AWS and Snowflake's AML triage walkthrough shows a practical AI automation pattern: assemble evidence, produce a structured brief, and keep regulated decisions with humans.

6 min read
Abstract geometric sphere with connected paths representing AI agents evaluating enterprise IT incidents.
AI Development·

Enterprise IT agents just got a harder benchmark. The best models still missed half the incidents.

ITBench-AA shows why enterprise IT agents need scoped pilots, workflow receipts, eval datasets, approval gates, and human escalation before they touch production systems.

8 min read
Abstract geometric workflow gate representing controlled AI agent decisions.
Industry Insights·

Claude Opus 4.8 Makes Agent Honesty a Business Requirement

Claude Opus 4.8 is stronger, but the real business story is whether AI agents can admit uncertainty, catch mistakes, and preserve review points.

7 min read
Abstract sound waves flowing into workflow nodes for realtime voice automation.
Industry Insights·

OpenAI's new realtime voice models turn speech into a workflow interface

OpenAI's May 2026 realtime audio models make voice more useful for business workflows. Here is how to choose between live voice agents, live translation, and streaming transcription.

9 min read
Abstract illustration of connected workflow nodes moving through protected infrastructure layers without text.
AI Development·

Google's Agent Executor shows why AI agents need runtime infrastructure

Google's Agent Executor points to a practical shift: production AI agents need durable execution, isolation, state consistency, recovery, and audit trails.

7 min read
Abstract circular workflow illustration showing human review and AI feedback loops without text.
AI Development·

OpenAI's tax agents show why AI automation needs a feedback loop

OpenAI's Tax AI pilot with Codex is less a story about automated tax prep and more a lesson in production AI: agents improve when practitioner corrections become structured evidence, evals, and guarded releases.

8 min read
Microsoft's Computer-Using Agents Are GA. The Real Story Is Legacy Workflow Automation.
AI Development·

Microsoft's Computer-Using Agents Are GA. The Real Story Is Legacy Workflow Automation.

Microsoft's Copilot Studio computer-using agents make AI-driven UI automation generally available. For SMB teams, the opportunity is not letting agents roam across screens. It is using governed workflows to bridge legacy systems that lack usable APIs.

8 min read