openspec-pi: Spec-Driven Development, Built Into Your Agent
openspec-pi is a pi package that wires OpenSpec spec-driven development into your coding agent — auto-injecting specs into every session, adding a…
AI, XR, DevOps, graph theory, knowledge graphs, and digital sovereignty. Deep dives, tutorials, and analysis from across the GraphWiz knowledge base.
openspec-pi is a pi package that wires OpenSpec spec-driven development into your coding agent — auto-injecting specs into every session, adding a…
Two NVIDIA DGX Sparks, a bonded 400 Gbps interconnect, and DeepSeek-V4-Flash on-premises — a complete look at sovereign AI inference in 2026.
How Neo4j's new AgentMemory SDK for .NET uses knowledge graphs as the persistence substrate for AI agent memory — enabling structured recall,…
We applied the CaRE compute-aware evaluation protocol (arXiv:2607.24763) to our agent benchmark suites — Agent Skill Bench and AMBench.…
How a templated substrate architecture enables heterogeneous LLM agents to collaborate on knowledge work — research synthesis, literature reviews,…
The GCC steering committee has announced its official AI policy — governing AI-generated code contributions to the GNU Compiler Collection.…
How Google's Gemini API managed agents — now with 3.6 Flash support and lifecycle hooks — let you deploy persistent, stateful AI agents without managing…
How Kernel Forge uses a multi-agent harness to generate, compile, profile, and iteratively optimise CUDA kernels from natural language descriptions —…
How Amazon Bedrock's explicit prompt caching for GPT-5.6 Sol, Terra, and Luna models reduces inference costs by up to 70% and latency by 50% — and how…
How to build production-grade job queues on PostgreSQL that scale to millions of jobs — covering SKIP LOCKED, partial indexes, priority queues, the…
How to build an inference meta-monitoring system for Amazon SageMaker AI endpoints using Amazon QuickSight — monitoring not just model metrics but the…
GitHub's 2026 security hardening across npm and GitHub Actions — what changed, how the attacks work, and how to defend your software supply chain…
Production frameworks for governing autonomous agents: bounded autonomy, inter-agent security, jailbreak prevention, and compliance under the EU AI Act.
How 3B-30B parameter models are moving to edge devices — from smartphones to industrial systems — driven by NPU hardware, quantization advances, and…
Production patterns for tiered inference routing: how to reduce LLM costs by 50-100x while improving latency and privacy by routing 70-80% of queries to…
Nanbeige's 3B-parameter Looped Transformer increases effective model capacity without adding parameters by reusing layers in a recurrent loop.…
Mind Lab's Macaron-V1 uses a Mixture-of-LoRA architecture with separate specialists for chat, agent, coding, and UI tasks.…
A practical guide to fine-tuning open source LLMs in production — when to fine-tune, LoRA vs QLoRA vs full fine-tuning, tooling landscape, data quality…
Get notified when I publish new articles on AI infrastructure, DevOps, and XR development.