Lab journal
Research notes and product updates.
Each article cites the product, repository, benchmark, or primary source used for its claims.
Systems · September 16, 2026
What a Personal WireGuard VPN Costs on AWS
Measured prices for a self-hosted WireGuard VPN on AWS, why Lightsail beats EC2 for this workload, and the user data trap that boots an instance with no VPN on it.
Read note →
Evaluation · September 9, 2026
Testing LLM Applications Without Fooling Yourself
Most of an LLM app is ordinary deterministic code and should be tested as such. The model-dependent part needs evals, and evals have failure modes that produce confident, wrong numbers.
Read note →
AI security · September 9, 2026
AI Guardrails That Hold: Architecture Over Prompts
You cannot instruct a model into being safe, because instructions and data are the same tokens. The controls that survive contact with an attacker sit outside the model entirely.
Read note →
LLM engineering · September 9, 2026
Structured Outputs: Valid JSON Is Not Correct JSON
Constrained decoding guarantees your schema parses. It guarantees nothing about the values inside it, and schema shape changes answer quality more than most teams realise.
Read note →
LLM engineering · September 9, 2026
Multimodal LLM Integration: Costs and Failure Modes
Sending an image to a model takes about four lines. What breaks afterwards is token economics, silent misreads, and an injection surface most teams do not know they opened.
Read note →
LLM infrastructure · September 9, 2026
OpenRouter Alternatives: Pick the Right Category First
Most alternatives lists mix four incompatible products together. Sort them by what you are replacing, and the shortlist gets short.
Read note →
LLM infrastructure · September 9, 2026
Open Source LLM Gateways: Library or Proxy?
LiteLLM, Portkey, Bifrost and freelm all remove the hosted router from your request path. They are not the same shape, and the shape decides more than the feature list.
Read note →
Evaluation · September 9, 2026
Why Agent Memory Benchmarks Cannot Be Trusted
Every memory vendor publishes numbers where they win. What those runs leave uncontrolled, and the open harness we built to score our own system by the same bar.
Read note →
LLM infrastructure · September 9, 2026
OpenRouter Alternatives With a Real Free Tier
Six providers that serve models at no cost, what each is good for, the failure modes nobody mentions, and how to route across them without rewriting your client.
Read note →
Search research · September 9, 2026
When Page One Stops Paying: A Measured Case
Ranking first on a non-brand query returned 4.25%, not the textbook 28%. What 91 days of Search Console and Bing AI citation data showed, and what it changes.
Read note →
Agent engineering · September 9, 2026
Giving an AI Agent Write Access to Your Paper
Four ways to connect an AI model to a LaTeX project, what each one can actually destroy, and the questions worth asking before you approve a consent screen.
Read note →
Product research · August 29, 2026
LetX in 2026: Lexi, MCP, and Academic Pro
What LetX is after the 28 August release: a compiler-aware agent, a hosted MCP server, and two paid plans. Figures rechecked on September 25, 2026.
Read note →
LetX engineering · August 29, 2026
How Lexi Checks LaTeX Edits with the Compiler
Lexi reads the active project, edits source, compiles the result, returns the PDF or error log, and preserves project history.
Read note →
LetX engineering · August 29, 2026
LetX MCP: What an AI Client Can Actually Do
A tool-by-tool explanation of the LetX MCP server, its OAuth boundary, and how project editing differs from an ordinary chat attachment.
Read note →
LetX templates · August 29, 2026
How LetX Generates LaTeX Template Previews
LetX stores template source, metadata, a compiled PDF, and a first-page image. Additional checks flag common preview problems.
Read note →
Comparison · August 29, 2026
LetX vs Overleaf: A Sourced Decision for Research Teams
A dated comparison of free collaboration, AI features, history, Git, pricing, and institutional options in LetX and Overleaf.
Read note →
Evaluation · August 29, 2026
Context Heavy: What the LoCoMo Result Shows
Context Heavy led a bounded LoCoMo retrieval comparison, but the sample size and latency trade-off are part of the result.
Read note →
Systems · August 28, 2026
Why QuantumSketch Uses Durable Workflows
Prompt-to-video generation is a distributed job, not one model call. QuantumSketch uses Temporal to keep multi-minute renders recoverable.
Read note →
Accessible computing · August 27, 2026
Bagh Language: Programming Logic in Bangla
Bagh is a bilingual interpreter and browser IDE exploring what changes when beginners can learn programming logic in Bangla.
Read note →
Open source · August 26, 2026
freelm: Routing Across Free LLM Providers
freelm exposes six free-tier providers through one interface, with explicit controls for failover, quotas, streaming, and paid-model avoidance.
Read note →
Publication · August 25, 2026
Research Note: Detecting Faint Exoplanets
A concise record of the 2024 BRAC University research work co-authored by Shihab Shahriar Antor.
Read note →