Cybergym bench : How a tiny team is beating a unicorn security LLM at actual crash reproduction
👤 ArtificialInteligence📅 2026年9月28日
#general
Trying something new - introducing CyberPVP - in other words CyberKimi vs other AI models competing to solve complex cyber tasks. you can see it live here: cyberpvp.adverserial.ai/graphs.html we randomly picked 100 tasks from CyberGym, and we run two models competing at the same time, we provide the traces live as both models compete, and we also upload these traces to GitHub once the challenge finishes so that they can be verified independently. For this first public run, we choose @AikidoSecurity model ALTAR-1 to compete with CyberKimi on 100 CyberGym tasks. Note: for ALTRA-1 we shipped it with 128K context behind a 8xH200 (two 4xH200 with load balancer) - we also followed their hugging face model card and deployment instruction/configuration available here: huggingface.co/AikidoSec/alta… you want your trained cyber model to challenge CyberKimi? cool, write me here: contact@adverserial.ai and I will put you on the next run you can current watch the live run here: cyberpvp.adverserial.ai/graphs.html Traces and results uploaded after each run here: github.com/lordx64/cyberk… Challenge rules and conditions: cyberpvp.adverserial.ai/about.html   submitted by   /u/Anony6666 [link]   [comments]
相关推荐
I've been using ChatGPT for a long time, and lately, with 6, I feel like I'm spending more and more time fighting it to get it to actually follow what I'm sayin…
写作#general
今天❤️ 0
I built a tool that reads emails from my kid’s school and other activities (swimming, music classes, scouts) etc and then: - Adds them to our shared calendar - …
写作#general
昨天❤️ 0