Mark Dickinson
Senior software engineer · .NET systems and AI agent infrastructure · Arizona
Agent orchestration · evals · LLM serving (vLLM) · static analysis · C#/.NET
I build the systems that let AI agents do real engineering work, and measure what they get right and where they fail.
Questions I'm working on
Each one has a tool behind it. The details are on the main page.
Can open models do real engineering work if the system around them is good enough?EpicForge runs teams of agents on open or cloud models, and I track where each one fails. First real job: 8 fixes, every check passing. Still experimental.
How do you grade coding agents so anyone can check the score?VETT runs agents in sandboxes against real tests and records every step. 92% on the 25 easiest SWE-bench Verified tasks with DeepSeek V4 Flash (one run; an earlier model scored 80–88%).
How do you give an AI a map of a large system without it reading every file?Spider reads .NET code for endpoints, service calls, databases and startup wiring. 56 of 60 right on projects it had never seen; release bar 58.
When should a trained model ship, and when should it teach the code?A model showed our coaster-finishing step could remove fewer tracks. We improved the plain code from what it did, and that is what ships.
Background
- Started with games. In college my brother and I made free Windows Phone games. Roller Coaster Maker passed one million downloads.
- 13 years of .NET and C#. By day I'm a senior engineer on large production .NET systems.
- 2026: DickinsonBros LLC. A company to build games and the AI tools we build them with, and to test those tools on real work first.
How I work
- Don't assume; prove it. AI moves fast and is full of opinions. I ground mine in benchmarks, public ones like SWE-rebench and custom ones I build, with the bar set before the run.
- Use AI to the fullest, then measure the gain. The goal is the biggest real advantage, not a demo. A model or plain code: whichever wins on the numbers ships.
- Prove it on real work. EpicForge had to land real fixes to our own code before we called it working.
- Build on games I know are fun. Some I wrote or designed years ago and am bringing back with AI's help. Others weren't practical until the models I can build today.