npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

ctv-mcp-server

v0.1.1

Published

MCP Server for Veeva Clinical Trial Vault (ctv.veeva.com) - detail extraction, RAG export, notification, reporting and query

Readme

CTV MCP Server

Veeva Clinical Trial Vault (ctv.veeva.com) 的 MCP 接入层,提供六大通用能力:临床详情提取、RAG 知识库导入、信息推送、报告生成、临床查询、指定临床查询。

架构参考 chictr-mcp-server,但因 CTV 提供 GraphQL API,无需无头浏览器与反爬对抗,依赖从 6 个降到 2 个,单条详情耗时约 0.55 s。

完整可行性评估见 dev/FEASIBILITY.md。


xybct CLI(Python,零依赖)

除 MCP Server 外,本仓库提供独立的 Python CLI,用于本地底册维护 + FastGPT 知识库月度增量同步。 入口 main.py,核心库 scripts/ctv_sync.py,仅用标准库。

pip install -e .            # 装出 xybct 命令;或免安装直接 python3 main.py

# 一次性配置(写入 ~/.ctv-mcp/config.json,权限 600)
xybct config --base https://your.fastgpt --apikey openapi-xxx \
             --dataset <datasetId> --collection <collectionId>

xybct import CTV_results.csv    # 导入 CSV 底册(windows-1252 自动识别)
xybct detail --limit 200        # 回源 GraphQL 补齐纳排标准等全字段
xybct push                      # 全量推送(自动长文分片)
xybct sync                      # 【月度增量】探测 → diff → 删旧推新 → 等训练
xybct status                    # 底册 / 远端 / 最近同步记录
xybct verify                    # 逐条核对覆盖,零丢失校验

增量同步只对 CTV 侧真正变化的研究做回源与重推,实测能省掉 95%+ 的 embedding 开销。 详见 dev/INCREMENTAL_SYNC.md、 FastGPT 接入实测见 dev/FASTGPT_VEEVA_SETUP.md。

独立技能包

skills/ctv-fastgpt-sync/ 是上述能力的自包含技能包, 可脱离本仓库单独部署(自带业务代码 + 部署脚本 + 配置校验 + 三份参考文档):

cd skills/ctv-fastgpt-sync
./scripts/setup.sh                 # 零依赖,不建 venv 不装包
python3 scripts/check_config.py    # 分级校验 + CTV/FastGPT 连通性探测
./scripts/install.sh               # 可选:挂到 ~/.agents/skills

快速开始(MCP Server)

从 npm 安装(推荐)

npm install -g ctv-mcp-server      # 或免安装直接用 npx ctv-mcp-server

从源码构建

npm install
npm run build

# 1) 通道自检:确认 CTV 三条通道仍可用(6 项检查)
node dist/scripts/probe.js

# 2) 端到端跑通六大能力
node dist/scripts/e2e.js /path/to/CTV_results_2026-08-25.csv

索引自愈

studies 与全文索引 studies_fts 的一致性由数据库触发器保证(trg_studies_fts_ai/au/ad), 任何写入路径(TS 服务端、Python CLI、手工 SQL)都无法绕过:

  • 新增/修改研究时自动维护 FTS,不存在「数据写了但搜不到」
  • 升级到本版本后,首次打开数据库会自动对账并补齐历史缺行(只补不删,可反复执行)
  • 若使用外部脚本批量写入,也可手动对账:
CTV_DATA_DIR=~/.ctv-mcp python3 scripts/ctv_sync.py reindex

接入 MCP 客户端

全局安装后直接 npx 拉起:

{
  "mcpServers": {
    "ctv": {
      "command": "npx",
      "args": ["-y", "ctv-mcp-server"],
      "env": {
        "CTV_USER_AGENT": "your-app/1.0 (+contact: [email protected])",
        "CTV_RPS": "2"
      }
    }
  }
}

源码方式构建则指向本地入口:

{
  "mcpServers": {
    "ctv": {
      "command": "node",
      "args": ["/Users/you/Downloads/ctv-mcp-server/dist/index.js"],
      "env": {
        "CTV_USER_AGENT": "your-app/1.0 (+contact: [email protected])",
        "CTV_RPS": "2"
      }
    }
  }
}

环境变量

| 变量 | 默认值 | 说明 | | --- | --- | --- | | CTV_DATA_DIR | ~/.ctv-mcp | SQLite 索引与导出产物目录 | | CTV_DB_PATH | <dataDir>/ctv.db | 索引文件路径 | | CTV_USER_AGENT | 内置占位 UA | 建议改为你方标识 + 联系方式,便于对方识别流量来源 | | CTV_RPS | 2 | 请求节流(每秒请求数),礼貌抓取 | | CTV_CONCURRENCY | 3 | 最大并发 | | CTV_TIMEOUT_MS | 20000 | 单请求超时 | | CTV_REDACT_CONTACTS | true | 研究者姓名/电话/邮箱默认脱敏,设 false 才输出 |

数据来源声明:本包不附带任何临床数据,仅提供抓取与建库工具。首次使用需自行取数: 用 sync_sitemap 枚举 slug 池,或从 ctv.veeva.com 导出 CSV 后经 import_csv_export 导入。 请遵守 ctv.veeva.com/robots.txt,并注意研究者联系方式属于个人信息。


数据通道

| 通道 | 用途 | 说明 | | --- | --- | --- | | GraphQL POST /graphql | 详情提取 | studyProfile(utn\|nct\|slug) → 37 字段,无需认证 | | Sitemap | 全量发现 | 60 分片 ≈ 597,907 条研究 slug,robots 声明的合规通道 | | CSV 导出 | 冷启动建库 | 35 列,windows-1252 编码,仅用于建索引 |

重要:ctv.veeva.com/robots.txt 禁止 /study-search,因此检索走本地 FTS5 索引而非在线代理。先建库再检索。

CSV 有损警告:CTV 导出侧编码有损(如土耳其语 ı → ?)。CSV 只用于建索引,权威文本一律回源 GraphQL。导入器会自动检测并告警。


工具清单

建库

| 工具 | 说明 | | --- | --- | | import_csv_export | 导入 CTV 导出的 CSV,自动处理 windows-1252 编码与有损检测 | | sync_sitemap | 从官方 sitemap 枚举 slug 池,按分片增量(每片约 1 万条) | | backfill_details | 批量把索引升级为全字段详情 |

六大能力

| 能力 | 工具 | 说明 | | --- | --- | --- | | 临床查询 | search_studies | FTS5 全文(标题/摘要/适应症/关键词/纳排标准)+ 状态/分期/国家/日期结构化过滤 | | 指定临床查询 | get_study_detail | UTN / NCT / slug / 详情页 URL 四种入口自动识别 | | 详情提取 | get_study_detail backfill_details | 三级取数:内存 → SQLite → GraphQL 回源 | | RAG 导入 | export_rag | 语义分节导出,支持 markdown / jsonl / text | | 报告生成 | generate_report | 聚合分析,输出 markdown / html / json | | 信息推送 | create_watchlist run_watchlist list_watchlists get_change_digest | 订阅 + 轮询 diff,产出可直接推送的 Markdown 摘要 | | 运维 | get_index_stats | 索引统计、缓存状态、HTTP 指标 |


典型工作流

# 建库
import_csv_export({ file_path: "~/Downloads/CTV_results_2026-08-25.csv" })
  → 193 行入库,编码 windows-1252,19 条有损告警

# 检索
search_studies({ keyword: "pancreatic cancer", status: ["Recruiting"], limit: 10 })
  → 命中 104

# 取详情(含 CSV 缺失的纳排标准)
get_study_detail({ study_id: "UTN000277244" })

# 灌进知识库
export_rag({ keyword: "KRAS", format: "jsonl", limit: 50 })
  → 自动补齐详情后分块导出

# 出报告
generate_report({ condition: "Pancreatic Neoplasms", format: "html" })

# 订阅监控
create_watchlist({ name: "胰腺癌招募中", keyword: "pancreatic", status: ["Recruiting"] })
run_watchlist({ name: "胰腺癌招募中" })
  → 返回 digest_markdown,可直接投递飞书/企微

GraphQL 详情比 CSV 多出的关键字段

| 字段 | 价值 | | --- | --- | | eligibilityCriteria | 纳入/排除标准全文,患者匹配核心依据,RAG 最高价值字段 | | keywords | 检索召回关键,如 ["68Ga","Claudin 18.2"] | | conditions | MeSH 标准疾病词(CSV 只有原始标签),支撑跨库术语对齐 | | armGroups | 试验分组设计(组别/类型/描述/干预) | | locations[].status | 站点级招募状态 + 联系人 |

CSV 中 Result First Posted Date 与其 Type 列填充率为 0%,实际可用 33 列。


RAG 分块策略

按临床语义切 7 节,而非机械定长切分——避免把纳入标准第 3 条和排除标准第 1 条切进同一块造成检索误导:

overview               概览(标识/状态/分期/申办方/人群/时间线)
brief_summary          研究摘要
detailed_description   详细描述
eligibility_criteria   纳入与排除标准
arm_groups             试验分组设计
interventions          干预措施
locations              研究中心

每块重复元数据头,保证任一块被单独召回时仍可溯源到具体研究与 UTN。

产物格式:

  • markdown — 每研究一个 .md,适合 Dify / FastGPT 目录导入
  • jsonl — 每行一 chunk,适合向量库批量 upsert
  • text — 纯文本合并

隐私

CTV 详情包含研究者真实姓名、手机号、邮箱。RAG 导出与报告默认脱敏,需显式传 include_contacts: true 才保留。

个人信息一旦进入向量库很难清理,请谨慎开启。全局默认可用 CTV_REDACT_CONTACTS=false 关闭(不建议)。


环境变量

| 变量 | 默认 | 说明 | | --- | --- | --- | | CTV_DATA_DIR | ~/.ctv-mcp | 数据目录(SQLite + 导出产物) | | CTV_DB_PATH | $CTV_DATA_DIR/ctv.db | 索引库路径 | | CTV_RPS | 2 | 请求速率上限(QPS) | | CTV_CONCURRENCY | 3 | 最大并发 | | CTV_TIMEOUT_MS | 20000 | 单请求超时 | | CTV_MAX_RETRIES | 3 | 最大重试次数 | | CTV_DETAIL_TTL_MS | 86400000 | 详情缓存 TTL | | CTV_REDACT_CONTACTS | true | 联系人脱敏开关 | | CTV_USER_AGENT | 内置 | 建议设为含真实联系方式的标识 |


目录结构

src/
├── index.ts                    MCP 服务入口 + 13 个工具注册
├── config.ts                   配置与 robots 规则
├── types.ts                    领域模型(对齐 GraphQL studyProfile)
├── runtime/
│   ├── http-client.ts          令牌桶限流 + 退避重试 + robots 守卫
│   ├── errors.ts               结构化错误码
│   └── input-validation.ts     主键识别 + 参数校验
├── sources/
│   ├── graphql-client.ts       studyProfile 查询(37 字段选择集)
│   ├── csv-importer.ts         windows-1252 解码 + RFC4180 解析 + 有损检测
│   └── sitemap-crawler.ts      60 分片枚举
├── store/
│   └── repository.ts           SQLite + FTS5 索引、变更日志、订阅
├── services/
│   ├── detail.ts               能力1 详情提取(三级缓存 + 脱敏)
│   ├── rag-export.ts           能力2 RAG 分块导出
│   ├── notify.ts               能力3 订阅巡检 + 摘要渲染
│   └── report.ts               能力4 报告聚合与渲染
└── scripts/
    ├── probe.ts                通道自检(6 项)
    └── e2e.ts                  六大能力端到端验证

故障排查

上游变更时先跑 probe.js。 它会逐项定位是通道变更还是本地逻辑问题。

GraphQL 的 introspection 已被服务端关闭,字段集固化在 graphql-client.ts。若报 ValidationError,说明字段名变更 —— 从详情页 HTML 的 __NEXT_DATA__ 复核当前 schema:

curl -s 'https://ctv.veeva.com/study/<slug>' \
  | grep -o '<script id="__NEXT_DATA__"[^>]*>.*</script>'

已知坑:字段名是 sponsors,不是 collaborators。


许可与合规

MIT(代码)。数据版权归 Veeva 及各申办方所有。

生产使用前请阅读 ctv.veeva.com/terms,遵守 robots.txt,并设置含真实联系方式的 CTV_USER_AGENT。