Blogs

2026

GPT 6 LLM Testing: What the Benchmarks Mean for Enterprise DeliverySep 8, 2026

What the GPT-6 benchmarks mean for enterprise AI delivery.

LLM Testing of Anthropic's Hacker-Opus TestSep 4, 2026

How Anthropic's model performs on the hacker-opus test.

LLM Testing of Gemini 3.8 Flash: Scored on the AI Matic BenchSep 4, 2026

Benchmark results for Gemini 3.8 Flash on the AI-matic bench.

The Missing Layer Between Your Agent and ProductionAug 26, 2026

Harness engineering: boundaries, repair loops, and the verification gate that decides what ships.

The Part of Frontend the LLMs Can't DoAug 23, 2026

Skills, component libraries, shader kits, icon sets, inspiration galleries, and design craft.

AI Agent Evaluation for Self-Improving Agent LoopsAug 21, 2026

An evaluation framework for autonomous agents that iterate and improve on their own execution loops.

Claude Watermarking and Enterprise AI TeamsAug 20, 2026

What Claude's watermarking approach changes for enterprise AI teams and provenance verification.

A Practical Guide to Graph Engineering by GoMLAug 18, 2026

GoML's practical guide to graph engineering, knowledge graphs, and agentic traversal patterns.

LLM Testing of Grok 4.6Aug 14, 2026

Grok 4.6 benchmarked, framed as a cost-curve shift rather than a capability jump.

The Ratio of Quiet ThingsAug 14, 2026

A good bookmark is a promise you make to a future self who has different interests and less time.

I Made My API 3,400x Faster, Then I Felt BadAug 13, 2026

A read-through cache story about Memcached, Neon Postgres, and the seductive lie of a big number.

LLM Testing of Claude Opus 5Aug 10, 2026

Testing Claude Opus 5 as the first frontier model GoML calls enterprise-ready.

Nine MinutesAug 4, 2026

The spiral took four evenings. The figure, nine minutes.

Debugging as PrayerJul 25, 2026

Both begin the same way: something is wrong and I do not know what.

LLM Testing of OpenAI GPT-5.6Jul 14, 2026

Evaluating OpenAI's GPT-5.6 variants, Sol, Terra, and Luna across reasoning and latency.

Grok 4.5 (High) Model OverviewJul 10, 2026

An overview and internal evaluation of Grok 4.5 (high) performance and tooling integration.

Instructions for a StreetlampJul 6, 2026

Stand where they left you. Let the moths mistake you for the moon.

The Complete Guide to Sonnet 5Jul 2, 2026

GoML's complete guide to Sonnet 5 architectures, context utilization, and latency optimization.

Sakana AI Fugu Enables One API for Production AIJun 23, 2026

How Sakana AI's Fugu enables one unified API for smarter model routing and resilient production AI.

The Harness Is the ProductJun 3, 2026

Why the orchestration harness matters more than the raw model.

Rogue Agent Impact VisualizerMay 28, 2026

A visualizer and blast-radius calculator for detecting and mitigating rogue AI agent behaviors.

Cutting a 3-Hour AWS Investigation to 11 MinutesMay 12, 2026

Cutting a 3-hour AWS observability incident investigation down to 11 minutes with automated telemetry agents.

The Complete Guide to AWS DevOps AgentMay 11, 2026

GoML's complete guide to configuring, deploying, and hardening the AWS DevOps Agent.

2025