CSB — The Mechanistic Oracle
csb computes what a perturbation does to a body, organ by organ, from cited mechanism — and refuses, by name, whatever it cannot compute. The refusal is not the caveat; it is half the product.
Why it exists
Ask what a compound does to a body and most systems answer by lookup — this drug, these recorded effects. That answer does not exist for a compound nobody has dosed, it cannot say why, and it cannot tell you what it is missing. csb computes the answer from mechanism instead, and where the mechanism is not in hand it says so by name rather than filling the gap, because a confident wrong number costs more here than a refusal with an address.
How it computes
- A drug is
(physchem, target binding, PK), never a name. A perturbation routes through one shared PBPK field → target occupancy → an expression gate → a per-organ effect. Nobody types "myopathy"; it emerges because HMGCR is expressed in muscle. The test is blunt — could a lookup table have produced this answer? If yes, the mechanism was decoration. - One shared field, so one changed fact moves every organ at once. Vancomycin oral vs IV is a single bioavailability fact, and it collapses the field everywhere together: the heart drops below the MRSA MIC (no endocarditis cover) while the kidney goes quiet. The safety/efficacy split falls out of the route, not out of a table.
- Contrast, not magnitude. A magnitude is
NOT_CALIBRATEDunless a calibration test says otherwise. The deliverable is the ratio that survives the uncited constants. - A refusal is a result with an address. An organ that cannot be computed says why, by name —
expression_unknown,reserve_not_characterised,no_pbpk_compartment— never a silent0. "We know the steps; we need k_cat for reaction X" is the most useful sentence the platform can hand a scientist. - Two solver worlds, one contract. Mechanistic kinetics (ODE/PDE, PBPK, genome-scale flux balance) and physics-informed ML (PINNs, neural ODEs) implement the same model contract, so a disease can be solved classically and learned, then the two cross-checked.
- Every claim re-checkable offline. A caveat ships inside the JSON, not in a docstring nobody reads. Models are sha256-pinned under their real names, and the literature is a ledger of what was actually fetched.
The corpus it computes on
Kinetics cannot say what is connected to what, so underneath the oracle sits its data layer: 59,161,677 typed, weighted, directed couplings over 28 sources, molecule to clinic. csb reads it through one flag, csb effect --with-corpus, and the corpus never reaches back.
- Strength and sign stay orthogonal. Negative feedback (TSH↔T4) is strongly coupled and negatively signed. A probability model has nowhere to put the minus sign, so inhibition ends up encoded as a small number and one of biology's most reliable couplings reads as weak evidence. STRING supplies coupling unsigned, SIGNOR supplies sign without strength — joined on the same accessions, they fill both fields of one edge.
- The path is the product, not the edge. What connects X and Y, and through what? Off-target reasoning then needs no special machinery: an off-target is a binding whose target is not on the path to the indication. Paths are scored against a degree-preserving null, so a route that node degree alone explains is refused too.
- Every edge knows how freely it may travel. 34,181,718 open · 23,470,039 gated · 1,509,920 share-alike, and
export(OPEN)returns exactly the open-track count, measured independently. Nothing gated leaks, at 59M scale.
The stack built on it
csb is the mechanistic oracle at the root — every cited constant enters here, and only here. Two regulatory/simulation surfaces sit on top of it.
the corpus — 59,161,677 typed couplings
over 28 sources. csb's data layer, read
through --with-corpus; it never reads back
|
v
csb
the mechanistic oracle — one shared
PBPK field; occupancy × expression
per organ, and it refuses by name
|
+--------------+---------------+
simulates efficacy grounds regulation
| |
CTAP RAPL
(dossier generation) (MFDS/NIFDS retrieval)
GLPI-103 · TILA-278 · OBX-319 provenanced · never generates
RAPL — Korea MFDS Regulatory Intelligence → ~53 GB of hard-to-reproduce Korea MFDS/NIFDS guidance turned into fast, provenanced retrieval — what an agent needs to ground a regulatory answer in the actual source rather than a plausible guess — plus a regulatory-timeline planner. The boundary is hard: it retrieves and structures, the agent reasons, and it never emits a clinical decision.
CTAP — Clinical Trial Analytics Platform →
One mock-drug study.yaml becomes a complete, internally consistent regulatory product: a simulated CDISC dataset (SDTM → ADaM → define.xml), the biostatistics on top of it (MMRM and estimands, reference-based imputation, tipping-point sensitivity, multiplicity gatekeeping, survival analyses), and a full ICH/CTD dossier across Modules 1–5 plus eCTD — with four automated gates keeping every document number traceable back to source data. Efficacy is simulated by the csb engine at this article's core.
Three mock programs exercise that structure across different targets, indications and submission routes. Compounds and patients are invented; every figure is simulated.
| Dossier | Target · Indication | Regulatory shape |
|---|---|---|
| GLPI-103 → | GLP-1 / Apelin dual agonist · Type 2 Diabetes | Full CTD, Modules 1–5 · NDA |
| TILA-278 → | anti-TL1A × IL-22R bispecific · Ulcerative Colitis | Biologic · Phase 2b package · BLA |
| OBX-319 → | anti-CD19 × anti-CD20 bispecific · Systemic Lupus | Biologic · Phase 3 submission · BLA |
Mock / virtual drugs for portfolio use — never real patient data, never an actual regulatory submission.
What makes an answer usable — IR · RA · Operator
Computing an effect and being allowed to claim it are two different things. Three pieces sit between them, each a boundary rather than a feature — the concept only; the detail lives with simpl-agent →.
- The product IR — described once, judged everywhere. A product is described once — composition, claims, capabilities, computed biology, provenance — and that one description travels to every surface above it. A rule pack may only ask fixed questions and take fixed actions; arbitrary expressions are refused, not missing. That is what makes a verdict reproducible instead of a prompt result.
- The regulatory ledger — what a computation may not claim. Before a single study is commissioned, every requirement the law asks is sorted into what a computation can speak to and what only a wet lab can answer. Naming the unreachable ones is the point: csb's refuse-by-name rule, applied to the claim rather than the constant.
- The operator — one runtime, imported by none of it. The agent drives the oracle and the surfaces above it, and no bench depends on being driven — so each stays swappable, and a conclusion is re-run rather than re-asked.
Stated plainly
Nothing has been checked by anyone who knows the biology — the sanity checks (APOE→Alzheimer, statins→rhabdomyolysis) were chosen by the same people who wrote the code. Magnitudes are not calibrated: the trajectory is right and the level runs low, and the output says so. And the three-route comparison does not work, because routes end in HP: terms and MedDRA terms with zero overlap — a licence blocker, not a scientific one, and the name-matching shortcut is refused.
목적
약이 몸에 무엇을 하는지 묻는 대부분의 시스템은 조회로 답합니다. 이 약, 이미 기록된 이 효과들. 그 답은 아직 아무도 투여해 본 적 없는 물질에는 존재하지 않고, 왜 그런지도 무엇이 빠져 있는지도 말하지 못합니다. csb는 그 답을 기전에서 계산하고, 기전이 손에 없으면 빈칸을 메우는 대신 이름을 붙여 거절합니다. 이 영역에서는 확신에 찬 틀린 숫자가 주소가 붙은 거절보다 훨씬 비싸기 때문입니다.
계산 방식
- 약은 이름이 아니라
(물리화학 특성, 표적 결합, 약동학)입니다. 교란은 하나의 공유 PBPK 장 → 표적 점유 → 발현 게이트 → 장기별 효과로 흐릅니다. 아무도 "근병증"을 입력하지 않습니다. HMGCR이 근육에 발현되기 때문에 근육이 스스로 답합니다. 판정 기준은 단순합니다. 조회 테이블로도 같은 답이 나왔다면 그 기전은 장식이었습니다. - 장이 하나이므로 사실 하나가 모든 장기를 동시에 움직입니다. 반코마이신 경구와 정맥 투여의 차이는 생체이용률이라는 사실 하나이고, 그 하나가 모든 장기의 장을 함께 무너뜨립니다. 심장은 MRSA MIC 아래로 내려가 심내막염을 덮지 못하고 신장은 조용해집니다. 안전성과 유효성의 갈림이 표가 아니라 투여 경로에서 나옵니다.
- 크기가 아니라 대비입니다. 보정 시험이 확인해 주지 않으면 크기 값은
NOT_CALIBRATED입니다. 결과물은 미인용 상수를 견디고 살아남는 비율입니다. - 거절은 주소가 붙은 결과입니다. 계산되지 않는 장기는 조용한
0대신expression_unknown,reserve_not_characterised,no_pbpk_compartment처럼 이유를 이름으로 말합니다. "단계는 알고 있고 반응 X의 k_cat이 필요하다"가 이 플랫폼이 연구자에게 건넬 수 있는 가장 쓸모 있는 문장입니다. - 두 세계가 하나의 모델 계약 뒤에 있습니다. 기전 동역학(ODE/PDE·PBPK·유전체 규모 대사 플럭스)과 물리 정보 기반 ML(PINN·neural ODE)이 같은 계약을 구현하므로, 같은 질환을 고전적으로도 풀고 학습으로도 풀어 서로 대조할 수 있습니다.
- 모든 주장은 네트워크 없이 다시 검증됩니다. 단서는 아무도 읽지 않는 독스트링이 아니라 JSON 안에 실려 나옵니다. 모델은 실제 이름과 함께 sha256으로 고정되고, 문헌은 실제로 가져온 것만 기록합니다.
계산의 바탕이 되는 코퍼스
동역학은 무엇과 무엇이 이어져 있는지를 말해 주지 못합니다. 그래서 오라클 아래에 데이터 층이 있습니다. 28개 출처를 하나의 문법으로 통일한 5,916만 개의 타입·가중치·방향이 있는 결합이 분자에서 임상까지 이어집니다. csb는 csb effect --with-corpus 플래그 하나로 이것을 읽고, 코퍼스는 반대로 csb를 읽지 않습니다.
- 세기와 부호가 서로 직교합니다. TSH와 T4 같은 음성 되먹임은 강하게 결합되어 있으면서 부호는 음수입니다. 확률 모델에는 마이너스를 넣을 자리가 없어 억제가 '작은 수'로 적히고, 생물학에서 가장 신뢰할 만한 결합이 약한 증거처럼 보이게 됩니다. STRING은 세기를 주지만 부호가 없고 SIGNOR는 부호를 주지만 세기가 없으므로, 같은 접근번호로 이어 붙여 한 간선의 두 칸을 채웁니다.
- 제품은 간선이 아니라 경로입니다. 무엇과 무엇이 무엇을 거쳐 이어지는가. 그러면 off-target 추론에 별도 장치가 필요 없습니다. off-target이란 적응증으로 가는 경로 위에 없는 결합일 뿐입니다. 모든 경로는 차수 보존 귀무 모형으로 채점하므로, 노드 차수만으로 설명되는 경로는 여기서도 거절됩니다.
- 모든 간선은 자신이 얼마나 자유롭게 이동할 수 있는지 압니다. 공개 3,418만 · 제한 2,347만 · 동일조건 151만이고,
export(OPEN)은 독립적으로 센 공개 트랙 수와 정확히 같은 수를 돌려줍니다. 5,900만 규모에서 제한 데이터가 새지 않습니다.
그 위에 올라가는 스택
인용된 상수는 오직 csb로만 들어옵니다. 그 위에 규제·시뮬레이션 표면 둘이 올라갑니다. RAPL은 재현하기 어려운 한국 MFDS·NIFDS 규제 자료 약 53 GB를 출처가 분명한 빠른 검색으로 바꾸고 규제 타임라인을 계획합니다. 경계는 분명합니다. 검색과 구조화만 담당하고 임상 판단은 생성하지 않습니다. CTAP은 하나의 가상 임상시험 정의(study.yaml)에서 CDISC 데이터·생물통계·ICH/CTD 전체 모듈·eCTD까지 만들고, 문서의 모든 숫자가 원천 데이터까지 추적되는지 네 개의 자동 게이트로 확인합니다. 유효성은 csb 엔진이 시뮬레이션합니다. 그 안의 GLPI-103·TILA-278·OBX-319는 표적과 적응증, 제출 경로가 서로 달라 구조를 실제로 시험합니다. 셋 다 가상 약물이며 실제 환자 데이터도 실제 규제 제출도 아닙니다.
답을 쓸 수 있게 만드는 것 — IR · RA · 오퍼레이터
효과를 계산하는 것과 그 효과를 주장해도 되는 것은 다른 문제입니다. 그 사이에 세 가지가 있고 셋 다 기능이 아니라 경계입니다. 여기서는 개념만 두고, 자세한 내용은 **simpl-agent →**에 있습니다.
- 제품 IR — 한 번 기술하고 어디서나 판정합니다. 조성·주장·기능·계산된 생물학·출처를 한 번만 기술하면 그것이 위의 모든 표면으로 그대로 전달됩니다. 규칙 팩은 정해진 질문만 묻고 정해진 조치만 취하며, 임의 표현식은 빠뜨린 기능이 아니라 거절된 것입니다. 그래야 판정이 프롬프트 결과가 아니라 재현 가능한 결과가 됩니다.
- 규제 원장 — 계산이 주장할 수 없는 것. 시험을 하나라도 발주하기 전에, 법이 요구하는 항목을 계산이 답할 수 있는 것과 실험실에서만 답할 수 있는 것으로 갈라 둡니다. 닿을 수 없는 쪽을 이름으로 밝히는 것이 핵심입니다. 상수를 이름으로 거절하던 규칙을 상수가 아니라 주장에 적용한 것입니다.
- 오퍼레이터 — 하나의 런타임, 그러나 아무도 그것을 import 하지 않습니다. 에이전트가 오라클과 그 위의 표면들을 구동하지만 어떤 벤치도 구동당하는 쪽에 의존하지 않습니다. 각각은 독립적으로 교체 가능하고, 결론은 다시 물어보는 대신 다시 실행해서 확인합니다.
밝혀 두는 한계
생물학을 아는 사람의 검증은 아직 없었습니다. 온전성 확인(APOE→알츠하이머, 스타틴→횡문근융해)은 코드를 쓴 사람이 직접 고른 것이라 도메인 검증이 아닙니다. 크기 값은 보정되지 않았습니다. 궤적은 맞고 수준은 낮게 나오며, 출력이 그렇다고 말합니다. 3-경로 비교는 작동하지 않습니다. 경로가 HP: 용어로 끝나고 MedDRA 용어와 겹치는 부분이 전혀 없기 때문입니다. 막고 있는 것은 과학이 아니라 라이선스이고, 이름 매칭으로 얼버무리는 지름길은 거절했습니다.
Live / 바로가기: csb.simplicity-is-art.com · Status / 상태: Beta (베타) · Pattern / 원칙: contrast over magnitude · refuse by name · auditable offline · Note / 참고: code closed, showcase public (코드 비공개, 쇼케이스 공개)
No comments yet. Be the first to say something!