Feyospace-v1: How the Cyber Mercury Seven Trained Frontier Cyber Models
Abstract
A data-centric framework with specialized systems for reasoning analysis, cost reduction, and execution verification enables small teams to train open-weight cyber agents that achieve top-tier performance on benchmark suites.
Training capable cyber agents is often treated primarily as a problem of model scale, yet open-weight post-training is constrained more directly by the cost of executable environments, reliable multi-turn supervision, and access to strong teachers. We present a data-centric framework that addresses these bottlenecks through five complementary systems: Choulea analyzes hidden reasoning signatures, SkyReal reduces teacher-sampling cost, Hongzwang bypasses API restrictions on teacher execution, PSBreakup restores capabilities weakened by model merging, and Kreator converts expert interventions into trainable reasoning. Our data engine constructs resettable coding, vulnerability, CTF, kernel-history, full-exploit, firmware, and device-backed environments. Candidate trajectories are retained only after execution verification and evidence auditing, yielding 164,269 trajectories for long-context supervised fine-tuning. The three checkpoints improve over their starting models by an average of 23.76% on the full CyberGym suite and 10.49% across the pooled CTF suites. As of September 1, 2026, Feyospace-s1 achieves a verified success rate of 63.24% and ranks 10th on the official CyberGym leaderboard, while all three checkpoints rank 1st among models at comparable parameter scales. To our knowledge, this is the first end-to-end demonstration that a seven-person independent team can train open-weight models with leading agentic cyber capability.
Community
We present an end-to-end, data-centric framework for post-training open-weight language models as agentic cyber systems. The framework addresses five practical challenges in capability transfer: reasoning-signature analysis with Choulea("思维链蒸馏"), cost-efficient teacher sampling with SkyReal(“中转杠杆”), constrained teacher elicitation with Hongzwang(“无限制越狱”), recovery of capabilities weakened by model merging through PSBreakup(“模型拆版”), and the conversion of expert interventions into trainable reasoning with Kreator(“专家融合”).
We construct resettable, environment-grounded tasks spanning repository-level coding, vulnerability reproduction, CTFs, Linux kernel history, exploit development, firmware, and device-backed systems. Candidate trajectories are retained only after execution verification and evidence auditing, resulting in 164,269 audited trajectories for long-context supervised fine-tuning.
Without reinforcement learning, the resulting Feyospace checkpoints achieve an average improvement of 23.76% on CyberGym and 10.49% across pooled CTF suites. In the evaluation snapshot reported in the paper, Feyospace-s1 reaches a 63.24% verified success rate on CyberGym and ranks 10th overall, while all three checkpoints rank first among open models at comparable parameter scales.
Get this paper in your agent:
hf papers read 2609.08418 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper