
Ep 55 - There Is No Best Model - Open Weights Caught Up and the Benchmarks Prove It - Dam Secure
Om avsnittet
đïž Coffee, Chaos and ProdSec, Ep 55
Your AI code reviewer costs eleven dollars a pull request. A model from China does the same job for forty two cents. Nobody saw that landing in September.
This week Cameron and Kurt sit down with Patrick Collins and Simon Harloff of Dam Secure, who benchmark every model drop against planted vulnerabilities in real pull requests. Patrick admits he told Simon six months ago that open weights would never catch the frontier labs. Simon told him he was wrong. The data agreed with Simon.
They get into why the best model in the field still misses a third of planted bugs, why Python recall craters compared to TypeScript, what happened when they asked an agent to break out of its own sandbox, and the DIY harness bill that runs past four hundred thousand dollars a year when you pick wrong.
If you work in Product Security, Application Security, DevSecOps, Security Architecture, or Cybersecurity and you are about to sign a twelve month contract on a model that re-ranks every six weeks, listen first.
â New episodes every Wednesday.
Coffee, Chaos and ProdSec -> strong coffee, stronger opinions.
Fler avsnitt
Visa alla avsnitt av Coffee, Chaos and ProdSecCoffee, Chaos and ProdSec med Cameron Walters and Kurt Hendle finns tillgÀnglig pÄ flera plattformar. Informationen pÄ denna sida kommer frÄn offentliga podd-flöden.