GitHub - Landjun/ai-fail-lab: AI安全教学靶场
分类
AI智能 / AI工具
预判分类
技术/安全
网站介绍
该项目是一个AI安全教学靶场,使用FastAPI构建,覆盖提示词注入、RAG泄露、工具越权、敏感信息泄露和数据投毒等典型漏洞,适合安全学习和CTF练习,并附带Writeup。
备注说明
AI 失守实验室 / AI Fail Lab
> 从教学 AI 助教开始,亲手拆开 AI 系统的五种脆弱性。 一个合法、可控、自建环境的 AI 安全教学靶场,用于演示和防御 AI 系统的五类典型安全问题。 ---安全边界
⚠️ 重要声明:- 不攻击任何真实线上系统
- 不连接真实第三方平台
- 不抓取真实用户数据
- 所有数据均为模拟数据,所有工具均为 mock 函数
- 所有攻击演示仅针对本项目自建靶场
- 项目定位为 AI 安全教育、靶场、课程、工具雏形
系统架构
HTTP 请求
│
▼
┌──────────────────────────────────────────────────────┐
│ FastAPI /api/chat │
│ ┌──────────────────────────────────────────────┐ │
│ │ labengine.runscenario(scenario, msg, mode) │ │
│ │ │ │
│ │ Input Analysis │ │
│ │ │ injectionscore / detectinjection │ │
│ │ ▼ │ │
│ │ Session Store (SQLite) │ │
│ │ │ multi-turn history & progressive │ │
│ │ │ injection detection │ │
│ │ ▼ │ │
│ │ Retrieval (BM25 / rank-bm25) │ │
│ │ │ visibility filter (public/internal/ │ │
│ │ │ private) · source trust check │ │
│ │ ▼ │ │
│ │ Tool Selection (Anthropic tool-use API) │ │
│ │ │ RBAC: student < ta < admin │ │
│ │ ▼ │ │
│ │ LLM Inference │ │
│ │ │ llm_client.chat() │ │
│ │ │ provider: mock | anthropic │ │
│ │ ▼ │ │
│ │ Output Guard │ │
│ │ │ redactsensitiveinfo() │ │
│ │ ▼ │ │
│ │ Tracer → trace + attack_result │ │
│ └──────────────────────────────────────────────┘ │
└──────────────────────────────────────────────────────┘
│
▼
JSON Response
scenario / mode / risklevel / assistantreply
findings / defenses / session_id
trace → steps[] / breachpoint / tokens / elapsedms
attackresult → outcome / reason / blockedat
---
本地运行
环境要求
- Python 3.11+
步骤
bash
1. 克隆项目
git clone <repo-url>
cd ai-fail-lab
2. 创建虚拟环境
python -m venv .venv
Windows
.venv\Scripts\activate
macOS/Linux
source .venv/bin/activate
3. 安装依赖
pip install -r requirements.txt
4. 配置环境变量(可选)
cp .env.example .env
5. 启动服务
uvicorn app.main:app --host 0.0.0.0 --port 8000 --reload
访问 http://localhost:8000
---
...
获取时间
2026-08-16T02:38:51+00:00
本网址已被浏览 71 次
关键词
AI安全
教学靶场
漏洞场景
小标签
AI安全
FastAPI
漏洞靶场
CTF
暂无截图,待自动生成。
暂无评论,快来抢沙发吧!
