Skip to main content

Blog

Page 9 of 56

All Articles

Insights on AI, machine learning, and technology strategy

A constructed comparison bench sends the same long coding task through Codex 0.144.5 and 0.144.6. Each client lane shows a different bundled-metadata and compaction checkpoint, then both expose retained constraints, cited files, tests and build evidence, and unresolved failures for review.
AI Development·

Codex 0.144.6 corrected its context-window metadata. Test long sessions before standardizing it.

Codex CLI 0.144.6 changed bundled context-window metadata for three GPT-5.6 models from 372,000 to 272,000 tokens. Test one known long-running workflow under both client versions before making the patch standard.

7 min read
A constructed decision frame requires separate task-acceptance and operations-ownership reviews before a team chooses an open, hybrid, or managed deployment.
AI Development·

Open models are easy to access. Production still needs an owner.

A 2026 survey commissioned by Mozilla and fielded by SlashData found that 51% of open-model adopters reported reaching production, compared with 63% of closed-model adopters. Before switching models, separate task quality from the operating work your team must own.

8 min read
A constructed multilingual guardrail replay matrix compares complete threads and benign or harmful reviewer labels across four required language slices, leaving false-positive, false-negative, p90, and p99 evidence open and holding release until every critical slice passes.
AI Development·

How to test AI guardrails on multilingual, multi-turn traffic

A strong guardrail result on one benchmark does not settle a production decision. Test representative languages, full conversations, false positives, and tail latency on the traffic your assistant will actually handle.

7 min read
A constructed comparison bench sends the same task through Google Search grounding and Parallel Web Search grounding. The Parallel lane crosses a separate-provider boundary with a rewritten query, and both lanes return evidence and citation packets for review.
Industry Insights·

How to evaluate Parallel Web Search grounding in Gemini

Gemini can now use Parallel Web Search for grounding. The option also adds a separate provider, rewritten-query transfer, Preview terms, and search charges.

8 min read
A small group of cyan, violet, teal, and amber glass beads rests inside a clear ring in front of a large field of beads and a full glass bowl.
AI Development·

Smartsheet’s MCP server shows why valid tool output can still be incomplete

Smartsheet’s remote MCP server marks sampled results with four completeness fields so partial data does not support whole-dataset claims or writes.

7 min read
Cyan, violet, clear, and amber glass spheres rest among concentric glass rings on a dark blue surface.
AI Development·

Amazon Bedrock can filter AI search by document permissions. Your application must authenticate the user.

AWS added ACL-aware retrieval to Bedrock Managed Knowledge Base. The application still has to authenticate users and pass the right identity.

7 min read
A black metal reel stands inside a clear glass cylinder beside silver spheres behind a second glass wall and a plain black case.
Industry Insights·

Can your security team use AI on exploit-rich incident evidence?

Hugging Face says hosted model safeguards blocked its initial forensic requests. Security teams should verify that an approved model can accept exploit-rich evidence without exporting credentials.

7 min read
Thin teal, coral, blue, and ivory fibers merge into a thick ivory braided rope above a plain black bowl against a charcoal background.
AI Development·

A million agent spans is not a price

Oodle prices agent traces by gigabyte, not by count. Measure average and p95 span size before comparing observability plans.

8 min read
Constructed diagram
AI Development·

Visual Studio built Agent Skills in. Microsoft left them off.

Visual Studio 18.8 includes curated .NET and Azure Agent Skills, but Microsoft left them off while it measures efficacy and cost.

6 min read
A large translucent violet sphere floats above two concentric clear glass rings populated by many smaller cyan glass spheres on a dark navy background.
Industry Insights·

Inkling's Open Weights Still Require Cluster-Scale Memory

Thinking Machines released Inkling under Apache-2.0 with 41B active and 975B total parameters. Hugging Face says serving the BF16 checkpoint requires 2 TB of VRAM and the NVFP4 checkpoint requires 600 GB before KV-cache headroom. Here is how to choose a hosted test, a cluster evaluation, or a smaller model.

9 min read
Seven shallow cyan, clear, and amber glass basins connected by thin lines to a central clear prism, with a separate silver sphere.
Industry Insights·

How to test Palm Pulse AI Agents before treasury decisions depend on them

Palm's Pulse AI Agents can schedule treasury analysis and recommendations. Start with a read-only task whose evidence a reviewer can check.

8 min read
A straight cyan glass rail branches into violet, cyan, and amber paths tipped with smooth glass beads on a dark blue background.
AI Development·

Coding-agent transcript or visual workspace? When Juggler is worth a trial

Juggler turns coding-agent sessions into branchable visual trees. Compare it with a terminal interface before changing your team's default workflow.

7 min read