📢gitzw.com上线了,功能陆续更新中,如有问题或反馈请在下方反馈/建议中给我们留言。

← News

安全★★★arXiv · 2026-07-16

Cost-Aware Evaluation of Offensive and Defensive Security Agents

Traditional security agent evaluations focus on peak offensive capabilities with generous resources but overlook operational costs. This paper assesses language-model security agents using a cost-success framework on offensive Cybench and defensive Splunk BOTS v1 challenges.

📌 Key points
  • Traditional evaluations emphasize high-budget offensive capabilities
  • Each reasoning step incurs operational costs
  • Models are compared at fixed cost in this study

📄 阅读原文(arXiv)↗

本页为 gitzw.com 基于公开来源的 AI 中文解读,非原文转载。

🔗 延伸到本站

📘agents 中文教程查看本站中文教程 →

📘 Latest tutorials

View tutorials →
wigolo:本地AI助手的智能搜索引擎
入门 · 7 章
30 分钟掌握 Voicebox:本地 AI 语音工作室
入门 · 7 章
用 awesome-mcp-servers 构建智能交互平台
入门 · 7 章
WinGet 快速指南:包管理新体验
入门 · 7 章

📦 Newly indexed

View ranking →
ai-agent-book★13.9k
《深入理解 AI Agent:设计原理与工程实践》(李博杰 著)开源主仓库:全书正文、编译版 PDF 与按章配套
ai-job-search★19.7k
AI-powered job application framework built on Claude Cod
i-have-adhd★6.3k
A skill for your coding agent to stop it from burying th
openship★5.8k
Self-hosted deployment platform
design.md★22.6k
A format specification for describing a visual identity
hallmark★12.3k
Anti-AI-slop design skill for Claude Code, Cursor, and C

📰 相关资讯

GhostLock, a stack-UAF that has existed in all Linux distributions for 15 yearsGitLost: We Tricked GitHub's AI Agent into Leaking Private ReposJanuscape: Guest-to-Host Escape in KVM/x86 [CVE-2026-53359]EU Council forces Chat Control via fast-trackEspionage Against the European ParliamentGoogle Loses Appeal Over Record $4.7B EU Antitrust Fine