문제. 범용 AI 는 항상 답합니다 — 근거가 없어도. 제약·의학에서 그럴듯한
오답 하나의 비용은 답을 못 얻는 것과 비교가 되지 않습니다. 이 서비스의 설계 원칙은 하나입니다:
지킬 수 있는 답만 한다. 모르면 모른다고, 문헌이 엇갈리면 엇갈린다고, 철회된 논문은
철회됐다고 말합니다.
① 신뢰성 — 전부 측정으로 세웠고, 분모를 함께 공개합니다
장치측정 결과 (분모 병기)왜 중요한가
보류 게이트AUROC 0.756 (PubMedQA 300문항·사전등록) — Claude Haiku(0.742) 초과. 답 없는 질문 0/5 통과·확신 0.012 보류틀릴 것 같으면 답하지 않음
판정 시 정확도~79% (yes/no 확정판정 · 서로 다른 두 근거 조건에서 78.6% / 79.2%로 재현)확정할 때는 맞음 — 조건이 바뀌어도 흔들리지 않음
유령 인용 차단답변의 모든 PMID 를 회수 목록과 대조 — 목록 밖 인용은 ⚠표식 (실사고 1건에서 도입)범용 AI 의 대표 사고인 지어낸 인용이 구조적으로 불가
철회 논문 감지Europe PMC 철회 플래그 전수 검사 — 철회 논문엔 🔴 대문짝 표식철회 논문 인용은 규제 문서에서 치명상
판단의 한계 명시지지/반박 자동 라벨을 검증(186줄·외부판사)했더니 73.5%라 스스로 철회 — 판정 대신 논문별 소견 서술로 후퇴못 하는 것을 안 하는 것이 신뢰의 실체
🔴 검증 방식 — 시험마다 실행 전에 성공/실패 기준을 커밋(사전등록 12건)하고,
채점은 자사 모델이 아닌 외부 판사(Claude)가 하며, 채점기 자체를 변이시험으로 먼저 검증했습니다.
기준 미달은 미달로 공개합니다 — 이 페이지의 철회 이력이 그 증거입니다.
검색 대상 — 무엇을 몇 건이나 뒤지는가
소스규모 (실측일 2026-09-03 · 각 기관 API 집계)성격
Europe PMC4,882만 문헌 (PubMed/MEDLINE 포함) · 그중 806만 편은 전문(full text)까지 열어 읽음동료심사 문헌·프리프린트
ClinicalTrials.gov60.1만 임상시험 — 실패·중단 시험의 사유 포함등록부 (논문에 없는 것)
FDA 라벨26.2만 공식 라벨 (박스경고·금기·용법)규제 문서
FDA FAERS2,069만 부작용 신고 집계 (신고≠인과 명시)시판 후 실사용 신호
Semantic Scholar2.2억+ 논문 — 고인용 지형·TLDR 요약 (정식 API 키)인용 그래프 + 요약
OpenAlex3.2억 학술 저작 — 인용 지형(고인용 논문) 파악용인용 그래프
UniProt인간 단백질 2만+ 검토(SwissProt) — 표적 기능·질환 연관 정본표적 큐레이션
ClinVar유전 변이 300만+ 임상적 의미 분류정밀의료 정본
FDA 회수 이력시판 후 회수(recall) 조치 — 문헌에 없는 것집행 기록
NIH RePORTER연구비 과제 수백만 건 — 논문보다 2~3년 앞선 신호연구 지형
PubChem화합물 공식 약리 문서 (작용기전)큐레이션 DB
웹 (opt-in)Brave 검색 + 페이지 본문 — 기본 꺼짐비동료심사 · 최신
🔴 위 수치는 이 페이지가 쓰는 각 기관의 공개 API 로 실측한 값이며(집계일 명기),
질문마다 이 전체에서 관련 문헌을 골라 논문 8~12편·시험·라벨을 근거로 좁혀 답합니다.
② 필요성 — 범용 AI·일반 검색이 못 하는 것
필요범용 AI / 웹검색이 서비스
실패한 임상시험안 보임 — 실패 시험은 논문이 안 나옴등록부에서 중단 사유 원문까지 (실측: "환자 순응도 불량으로 중단")
규제 공식 문서웹 요약에 섞여 출처 불명FDA 박스경고·금기 원문 + FAERS 신고 집계 (신고≠인과 명시)
논문 전문초록·스니펫만오픈액세스 전문을 열어 읽음 — 실측: 요약만 77.5% → 전문 84.5% (분모 200)
데이터 주권질문이 외부 API 로 나감온프렘 — 9B 모델이 자체 GPU 에서. 웹검색은 기밀질의 금지 표식과 함께 opt-in
근거 등급 구분없음동료심사/프리프린트/등록부/규제문서/웹을 모두 다른 표식으로
③ 가치 — 왜 9B 로 충분한가 (실측)
같은 200문항에서 3배 큰 26B 가 9B 보다 1%p 낮았습니다. 근거를 손에 쥐여주면
9B 가 그중 93%를 정확히 읽습니다 — 병목은 모델 크기가 아니라 근거 회수였고, 그것을 7종 소스
(문헌 전문·리뷰 슬롯·임상등록부·FDA·작용기전·인용지형·웹 opt-in)로 풀었습니다.
작은 모델이라 회사 서버 한 대에 들어가고, 그래서 온프렘 납품이 성립합니다. 응답 11초.
🔴 정직한 한계 — ① 범용 지식·코드·긴 대화는 범용 모델이 낫습니다(대체재가
아니라 다른 도구입니다). ② "효과가 있는가" 류 판단형 질의엔 결론을 강제하지 않고 논문별 소견을
나열합니다 — 판단은 연구자의 몫이고, 저희가 검증한 것은 그 소견의 출처입니다. ③ 위 수치의
분모(200~300문항)는 파일럿 규모입니다. 확대 검증은 진행 중이며, 분모 없이 수치를 말하지 않습니다.
The problem. General-purpose AI always answers — even without evidence.
In pharma, one plausible wrong answer costs far more than no answer. Our design principle:
only give answers we can stand behind. If we don't know, we say so; if the literature
is split, we show the split; if a paper is retracted, we say retracted.
① Reliability — everything below was measured, denominators included
MechanismMeasured resultWhy it matters
Abstention gateAUROC 0.756 (PubMedQA 300 Qs, pre-registered) — above Claude Haiku (0.742). Unanswerable questions: 0/5 passed, conf 0.012When unsure, it does not answer
Accuracy when it commits~79% (replicated at 78.6%/79.2% under two different evidence conditions)Stable across conditions
Ghost-citation guardEvery cited PMID checked against retrieved evidence; unknown ones get a ⚠ flag, real-but-unretrieved ones are verified against Europe PMCFabricated citations are structurally blocked
Retraction detectionEurope PMC retraction flags checked on every paperCiting retracted work is fatal in regulatory contexts
Self-retraction recordWe validated our own support/refute labels (186 lines, external judge), got 73.5%, and removed the featureNot doing what we can't do — that is what trust is
🔴 Method — every experiment's pass/fail criteria are committed before running
(12 pre-registrations); grading is done by an external judge (Claude), and the graders
themselves are mutation-tested first. Failures are published as failures.
What we search — sources and sizes (live-measured 2026-09-03)
SourceSizeKind
Europe PMC48.8M records; 8.06M full-text read in fullPeer-reviewed + preprints
ClinicalTrials.gov601K trials incl. failed/terminated with reasonsRegistry
FDA labels / FAERS262K labels · 20.7M adverse-event reportsRegulatory
UniProt / ClinVar / PubChemCurated target function, 3M+ variant classifications, mechanism of actionCurated reference
Semantic Scholar / OpenAlex / NIH RePORTER220M+ / 322M works · grant landscapeCitation graph
Web (opt-in)Brave search + page reading — off by defaultNon-peer-reviewed
② Why general AI / web search can't do this
NeedGeneral AIThis service
Failed trialsInvisible — failures don't get papersRegistry termination reasons verbatim
Regulatory documentsMixed into web summariesFDA boxed warnings & contraindications verbatim + FAERS
Full textAbstracts/snippetsOpen-access full text read — measured 77.5%→84.5% (n=200)
Data sovereigntyQuestions leave to external APIsOn-premise 9B model on our GPU; web search is opt-in with a warning
③ Why a 9B model is enough (measured)
On the same 200 questions, a 3× larger 26B scored 1pp lower. Given the
evidence, the 9B reads 93% of it correctly — the bottleneck was retrieval, not model size,
and retrieval is solved with the 12-source stack above. Small enough for a single
on-prem server; ~11s per answer.
🔴 Honest limits — ① General knowledge, code, long chat: use a general
model; this is a different tool, not a substitute. ② For judgment questions ("is it
effective?") we do not force a yes/no — we list per-paper findings and leave the judgment
to the researcher. ③ The denominators above (200–300 questions) are pilot-scale; we never
quote numbers without them. Not a substitute for clinical judgment.