Hello! My name is Matt Suiche. I work on AI Security at Tolmo, and I also experiment with side projects in AI Safety (Weightless, etc.) and Emulation & Operating System Research (WASM PSX, WASM NanoKrnl, etc.). I recently discussed cyberwar in the age of AI, Iran’s cyber capabilities, and how AI is reshaping hacking on Bloomberg’s Odd Lots and the National Security Lab podcast.
Previously, I founded OnDB Inc., a data infrastructure startup for the agentic economy, and co-founded CloudVolumes (acquired by VMware in 2014) and Comae Technologies (acquired by Magnet Forensics in 2022), where I later served as Head of Detection Engineering. I also founded the cybersecurity community project OPCDE.
My path into technology started in reverse engineering as a teenager, and has since spanned memory forensics, operating systems, virtualization, blockchain, and now AI infrastructure.
Inkling-Small-NVFP4 now serves on 2x DGX Spark (TP=2) on stock vLLM v0.28.0 with CUDA graphs on: 78.3 GiB of model, 105k tokens of KV, no custom image, no eager mode. It took two patches: a Triton/SDPA rel-attention fallback for sm_121 (FA4 …
Follow-up to 'Abliteration Without Redistributing the Model': steering seven frontier open-weight models with runtime control vectors. Refusal is not equally sticky across model families (Qwen folds at alpha=1.0, the GLM-5.3 flagship gets …
Every abliterated model on HuggingFace is a full re-upload: 30.9 GB for orcarouter's Qwen3.8-27B, 166.9 GB for Keys' DeepSeek V4 Flash. The same change fits in 8.6 MB and runs on stock tooling. Weight editing, LoRA and runtime projection …