01
Multimodality
Vision-language systems that judge content, not just caption it.
- Visual-Qwen↗ · CLIP + Q-Former + Qwen3 4B, 92% eval accuracy, trained on an H200
- MicroMARC · vision-language model that flags cognitively degrading short-form video
Independent research across intelligence, compute, and data. We run the tests ourselves and publish the code, methods, and raw numbers.
Methodology and raw logs: snapdragon-vs-m5, snapdragon-vs-mediatek
Vision-language systems that judge content, not just caption it.
One PyTorch workload, profiled across GPUs, TPUs, NPUs, and phones.
Datasets and pipelines built to be rerun, not just cited.
No vendor allegiance. We buy or rent the hardware we test.
Methods, code, and raw logs ship with every result.
Claims come from benchmarks we ran, not spec sheets.