- Security research
- Reward signals
- Objective verification
Project 07Research preview
malagent
Exploring RLVR for security research.
malagent.ioResearch focus.
Applies Reinforcement Learning from Verifier Rewards to security domains where outcomes can be objectively verified. Code compiles, tests pass, detections fire. Configurable reward signals from binary pass/fail to graduated rewards keyed on detection severity, with verification modes from compile-only to full EDR.