This is the full developer documentation for PoktaCare Docs
# Inicio
> Hub técnico del estudio Grader — arquitectura, onboarding y enlaces a cada repo involucrado.
Este sitio es el hub técnico para quienes colaboran en el **estudio comparativo de extracción clínica** y su producto central, el **Grader**. Si acabas de unirte, empieza en [Primeros pasos](/getting-started/).
Acceso restringido: si estás viendo esto, ya tienes acceso autorizado. El contenido aquí es tan sensible como los repos privados que enlaza — trátalo igual, no lo reenvíes ni lo publiques fuera de este grupo. Ver [Reglas](/guardrails/) para el detalle.
**Agentes y LLMs:** empieza en [`/llms.txt`](/llms.txt). Usa [`/llms-small.txt`](/llms-small.txt) para contexto reducido o [`/llms-full.txt`](/llms-full.txt) para el sitio completo.
## El proyecto, en corto
[Sección titulada «El proyecto, en corto»](#el-proyecto-en-corto)
Automatizamos la preparación de registros clínicos para un registro de biológicos en reumatología (nombre del registro y de las entidades colaboradoras en modo *stealth* — ver [Reglas](/guardrails/)). Los médicos tomaban su nota de consulta y tenían que volver a capturarla manualmente en la plataforma del registro; el pago por hacerlo bajó de \~$30 a \~$12 USD por paciente y varios médicos ya no lo hacen. La apuesta: un motor que lea la nota clínica, prepare un borrador estructurado y ayude al médico a corregirlo y aprobarlo antes de cualquier envío controlado.
Para justificar esa apuesta con números corre un estudio comparativo — varios “harnesses” (arneses de extracción) sobre el mismo corpus de notas reales:
1. **LLM plano** — sin prompt especializado, línea base. Vive en `pokta-care-monorepo`, no es un servicio separado.
2. **RheumaAI** — agente ya entrenado en reumatología. Repo propio, servicio propio.
3. **Grader** — el producto en evolución. Repo propio, servicio propio, fork de RheumaAI.
4. Un cuarto brazo opcional (LLM ajustado por otro reumatólogo) — registrado en código como placeholder, no corre todavía.
Cómo se conectan estas piezas — quién llama a quién, dónde caen los resultados — está en [Arquitectura](/architecture/), no aquí; esta página es solo el mapa de “qué es qué.”
**North Star:** el objetivo final no es “leer un Word y llenar un formulario” — es un copiloto en tiempo real que, durante la consulta, va extrayendo la información clínica en vivo y le dice al médico “te faltó preguntar X” antes de que termine. La extracción retrospectiva de notas es el entrenamiento y validación de ese motor, no el producto final.
## Los repos
[Sección titulada «Los repos»](#los-repos)
| Repo | Qué es | Servicio |
| ----------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------- | ------------------------------- |
| [`pokta-grader-bioagent`](https://github.com/poktalabs/pokta-grader-bioagent) | El Grader — arm 3 del estudio | Railway, `grader.poktacare.com` |
| [`rheuma-ai-bioagent`](https://github.com/poktalabs/rheuma-ai-bioagent) | RheumaAI — arm 2 del estudio | Railway, `rheumai.xyz` |
| [`pokta-care-monorepo`](https://github.com/poktalabs/pokta-care-monorepo) | Orquestador del estudio: arm 1, scoring, persistencia, y la app de producto (`app.poktacare.com`) | Vercel + Railway |
Detalle completo, con links directos a README/CONTRIBUTING de cada uno, en [Repos y guías](/repos/).
## Equipo
[Sección titulada «Equipo»](#equipo)
* **Dr. Erick Zamora Tehozol** — Reumatología y validación clínica
* **Ing. Ángel Meléndez Córdoba** — Ingeniería
## Dónde seguir
[Sección titulada «Dónde seguir»](#dónde-seguir)
* **[Primeros pasos](/getting-started/)** — checklist de acceso, por dónde empezar según en qué repo vas a trabajar.
* **[BiobadamexAI](/biobadamexai/)** — workflow actual: ingesta, extracción, revisión clínica y envío controlado.
* **[Arquitectura](/architecture/)** — cómo se conectan los tres repos: quién llama a quién, cómo se mide, dónde caen los resultados.
* **[Repos y guías](/repos/)** — directorio de enlaces.
* **[Reglas](/guardrails/)** — las reglas que no se negocian. Léelas antes de tu primer cambio.
# BiobadamexAI
> Cómo una nota clínica pasa de la ingesta a extracción, revisión médica y registro controlado en BIOBADAMEX.
BiobadamexAI es el flujo institucional que convierte notas clínicas en borradores estructurados para BIOBADAMEX. Separa la automatización de la decisión médica: el sistema extrae y calcula; el médico revisa, corrige y aprueba; un paso posterior y explícito puede enviar el registro.
## En una frase
[Sección titulada «En una frase»](#en-una-frase)
```
flowchart LR
A["Nota clínica"] --> B["Inventario cifrado"]
B --> C["Extracción solicitada"]
C --> D["LLM + reglas deterministas"]
D --> E["Borrador por revisar"]
E --> F{"Decisión médica"}
F -->|corregir| E
F -->|rechazar| R["Rechazada"]
F -->|aprobar| G["Aprobada"]
G --> H{"Confirmar envío"}
H --> I["Cola serial"]
I --> J["Sandbox Tenki"]
J --> K["BIOBADAMEX crdA–crdE"]
K --> L["idpac + resumen"]
```
## Estados que ve el médico
[Sección titulada «Estados que ve el médico»](#estados-que-ve-el-médico)
| Estado | Significado | Acción disponible |
| ------------------- | --------------------------------------------------------------- | ---------------------------------- |
| Sin extraer | El archivo está guardado, pero no se ha enviado al extractor | Extraer |
| En proceso | Hay un trabajo de extracción activo o en reintento | Esperar |
| Error de extracción | Se agotaron los intentos y existe un fallo visible | Reintentar |
| Por revisar | Existe un borrador estructurado | Revisar y corregir |
| Aprobada | El médico aprobó el borrador | Enviar mediante un paso separado |
| Rechazada | El médico descartó esa extracción | Conservar el historial o reextraer |
| Registrada | El trabajador obtuvo un `idpac` y confirmó la página de resumen | Consultar el recibo |
Estos estados se derivan de archivos, trabajos y borradores. No existe una columna única que pueda desincronizarse del flujo.
## Superficies actuales
[Sección titulada «Superficies actuales»](#superficies-actuales)
* Consola médica: carga, inventario, extracción manual, revisión, corrección, aprobación, rechazo y envío.
* Ingesta por servicio: carga por lotes autenticada con clave de servicio. Solo guarda archivos; no prueba extracción.
* Auditoría desde RheumAI: califica una nota sin persistirla o crea un borrador `review_pending` para revisión.
* Agente clínico dedicado: inventario y extracción limitados al médico configurado; no aprueba ni registra.
* Panel de flujos: proyección operativa con estados, tiempos, modelo, número de correcciones e `idpac`, sin devolver el cuerpo clínico.
## Reglas centrales
[Sección titulada «Reglas centrales»](#reglas-centrales)
1. Un archivo cargado no equivale a un registro extraído.
2. Un borrador extraído no equivale a un registro aprobado.
3. Aprobar no envía automáticamente.
4. Enviar requiere propiedad, estado aprobado, campos requeridos resueltos y una bandera de entorno explícita.
5. El valor `unknown` se conserva. Un cero real o `false` documentado no se confunde con ausencia.
6. Las comorbilidades no documentadas solo se convierten a `No` después de confirmación explícita del médico.
7. Nada de esto demuestra paridad completa con todos los controles de la plataforma externa.
## Alcance de esta documentación
[Sección titulada «Alcance de esta documentación»](#alcance-de-esta-documentación)
Las páginas describen el código de `poktalabs/pokta-care-monorepo` en el commit `b852f056d527d5bcec72b67256e5af98d386b323`. No inspeccionan variables del entorno desplegado, datos clínicos, credenciales ni la plataforma externa en vivo.
## Evidencia en código
[Sección titulada «Evidencia en código»](#evidencia-en-código)
* [Rutas montadas por la API](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/api/src/index.ts)
* [Inventario de notas](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/web/src/pages/BiobadamexNotes.tsx)
* [Contratos de trabajos](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/packages/biobadamex-registry/src/jobs.ts)
* [Estado derivado de una nota](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/api/src/services/biobadamex-note-status.ts)
# BiobadamexAI
> How a clinical note moves through intake, extraction, clinician review, and controlled BIOBADAMEX registration.
BiobadamexAI is the institutional workflow that turns clinical notes into structured BIOBADAMEX drafts. Automation and clinical decisions remain separate: software extracts and computes; a clinician reviews, corrects, and approves; a later explicit step may submit the record.
## In one sentence
[Section titled “In one sentence”](#in-one-sentence)
```
flowchart LR
A["Clinical note"] --> B["Encrypted inventory"]
B --> C["Extraction requested"]
C --> D["LLM + deterministic rules"]
D --> E["Pending-review draft"]
E --> F{"Clinician decision"}
F -->|correct| E
F -->|reject| R["Rejected"]
F -->|approve| G["Approved"]
G --> H{"Confirm submission"}
H --> I["Serialized queue"]
I --> J["Tenki sandbox"]
J --> K["BIOBADAMEX crdA–crdE"]
K --> L["idpac + summary"]
```
## Clinician-facing states
[Section titled “Clinician-facing states”](#clinician-facing-states)
| State | Meaning | Available action |
| ---------------- | ---------------------------------------------------- | ------------------------------ |
| Not extracted | File is stored but has not entered extraction | Extract |
| Processing | Extraction is active or retrying | Wait |
| Extraction error | Attempts were exhausted and a visible failure exists | Retry |
| Pending review | Structured draft exists | Review and correct |
| Approved | Clinician approved the draft | Submit through a separate step |
| Rejected | Clinician rejected that extraction | Preserve history or re-extract |
| Registered | Worker obtained `idpac` and confirmed summary | Inspect receipt |
These states are derived from uploads, jobs, and drafts rather than stored in one status column.
## Current surfaces
[Section titled “Current surfaces”](#current-surfaces)
* Clinical console for upload, inventory, extraction, review, correction, approval, rejection, and submission.
* Service-key batch intake. It stores notes and does not prove extraction.
* RheumAI audit surface that grades without persistence or creates a held review draft.
* Dedicated clinician agent limited to one configured clinician. It cannot approve or register.
* PHI-light lifecycle dashboard with stage, timing, model, correction count, and `idpac`.
## Central rules
[Section titled “Central rules”](#central-rules)
1. Upload does not equal extraction.
2. Extraction does not equal approval.
3. Approval does not auto-submit.
4. Submission requires ownership, approved state, resolved required fields, and explicit enablement.
5. `unknown` stays distinct from measured zero and documented `false`.
6. Undocumented comorbidities become `No` only after explicit clinician confirmation.
7. Current behavior does not prove full parity with every external registry control.
## Scope
[Section titled “Scope”](#scope)
These pages describe `poktalabs/pokta-care-monorepo` at commit `b852f056d527d5bcec72b67256e5af98d386b323`. They do not inspect deployed variables, clinical data, credentials, or the live external registry.
## Code evidence
[Section titled “Code evidence”](#code-evidence)
* [Mounted API routes](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/api/src/index.ts)
* [Notes inventory](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/web/src/pages/BiobadamexNotes.tsx)
* [Job contracts](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/packages/biobadamex-registry/src/jobs.ts)
* [Derived note state](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/api/src/services/biobadamex-note-status.ts)
# RheumAI
> Qué es RheumAI, cómo procesa una consulta clínica y dónde encaja con BiobadamexAI.
RheumAI es un sistema especializado de apoyo a decisiones clínicas en reumatología. Combina conversación clínica, recuperación de literatura, herramientas médicas y una revisión posterior de la respuesta. Su propósito es ayudar al médico a organizar información, detectar vacíos y revisar hipótesis. Esta descripción explica su función; no afirma validación prospectiva, aprobación regulatoria ni condición de dispositivo médico. No sustituye el juicio clínico ni aprueba registros de forma autónoma.
## En una frase
[Sección titulada «En una frase»](#en-una-frase)
RheumAI convierte una pregunta o nota clínica en una respuesta estructurada y revisada por varias capas de software:
```
flowchart LR
A["Pregunta o nota"] --> B["Planificador"]
B --> C["Fuentes y herramientas"]
C --> D["Respuesta clínica"]
D --> E["ORVS y control de citas"]
E --> F["Revisión ética"]
F --> G["Respuesta para el médico"]
```
## Superficies actuales
[Sección titulada «Superficies actuales»](#superficies-actuales)
| Superficie | Uso | Estado observado |
| -------------------------------------------------------- | ------------------------------------------------- | ----------------------------------------------------------------------- |
| Aplicación web | Conversación, archivos y seguimiento de sesiones | Activa en el código principal |
| `POST /api/chat` | Flujo conversacional persistente | Activo; usa base de datos y puede usar archivos |
| `POST /v1/study/rheumaai-reply` | Evaluación controlada de una nota desidentificada | Activo; stateless, autenticado y sin escritura en las tablas de RheumAI |
| Modos `vanilla`, `rag`, `dag`, `quick-orvs`, `full-orvs` | Comparar variantes del pipeline | Experimentales; se activan explícitamente |
| Investigación profunda | Búsqueda y síntesis de evidencia más larga | Activa como flujo separado |
| x402 | Cobro opcional para clientes de API | Integración opcional, controlada por configuración |
## Qué sucede con una consulta clínica
[Sección titulada «Qué sucede con una consulta clínica»](#qué-sucede-con-una-consulta-clínica)
1. La capa de acceso aplica autenticación, límites y, si está habilitado, el control x402.
2. RheumAI crea o recupera la conversación y prepara el estado de la solicitud.
3. Si hay archivos, los analiza antes de planear la respuesta.
4. El planificador elige herramientas y fuentes. Entre ellas puede haber PubMed, Semantic Scholar, conocimiento local, grafo de conocimiento o interpretación de laboratorio.
5. Las fuentes recuperadas se reordenan según su relación con la consulta.
6. El generador produce la respuesta con la persona clínica de RheumAI.
7. ORVS revisa la respuesta cuando corresponde. El sistema puede regenerarla si no pasa.
8. Un control separado revisa PMID recuperados y elimina identificadores no verificados.
9. La revisión ética puede bloquear una salida rechazada. Si la revisión falla técnicamente, el flujo actual deja pasar la respuesta y registra el fallo.
10. El resultado vuelve al médico con el contexto disponible en ese turno.
## Qué no hace todavía
[Sección titulada «Qué no hace todavía»](#qué-no-hace-todavía)
* No produce por sí solo un registro BIOBADAMEX estructurado y listo para enviar.
* No calcula una puntuación de completitud de la **nota de entrada** contra todos los campos obligatorios del registro.
* No distingue en un contrato estructurado entre dato ausente, no aplicable y desconocido.
* No reemplaza la revisión y aprobación del médico.
La integración propuesta con BiobadamexAI usa RheumAI como capa de razonamiento y explicación. El esquema del registro, la procedencia campo por campo, la evaluación de completitud y la aprobación clínica permanecen en el flujo institucional controlado. Consulta [Evaluación de notas y completitud](/rheumai/clinical-notes/) y el [workflow actual de BiobadamexAI](/biobadamexai/).
## Límites de interpretación
[Sección titulada «Límites de interpretación»](#límites-de-interpretación)
La documentación describe el código en el commit `149508b3ab79da0b88c56deb63f6fb66e35d0e68`. Que una herramienta esté registrada no demuestra que esté configurada en todos los entornos ni que haya sido validada clínicamente. Las rutas experimentales y los fallbacks están etiquetados en las páginas siguientes.
## Evidencia en código
[Sección titulada «Evidencia en código»](#evidencia-en-código)
* [Entrada del servidor y rutas](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/index.ts)
* [Flujo principal de chat](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/routes/chat.ts)
* [Registro dinámico de herramientas](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/tools/index.ts)
* [Ruta stateless para el estudio](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/study/rheumaai-reply-route.ts)
# RheumAI
> What RheumAI is, how it processes clinical questions, and how it fits with BiobadamexAI.
RheumAI is a specialized clinical decision-support system for rheumatology. It combines clinical conversation, literature retrieval, medical tools, and post-generation review. Its purpose is to help clinicians organize information, find gaps, and examine hypotheses. This description explains its function; it does not claim prospective validation, regulatory approval, or medical-device status. It does not replace clinical judgment or autonomously approve registry data.
## In one sentence
[Section titled “In one sentence”](#in-one-sentence)
RheumAI turns a clinical question or note into a structured response reviewed by several software layers:
```
flowchart LR
A["Question or note"] --> B["Planner"]
B --> C["Sources and tools"]
C --> D["Clinical response"]
D --> E["ORVS and citation controls"]
E --> F["Ethical review"]
F --> G["Response for the clinician"]
```
## Current surfaces
[Section titled “Current surfaces”](#current-surfaces)
| Surface | Use | Observed state |
| -------------------------------------------------- | --------------------------------------------- | ------------------------------------------------------------------------ |
| Web application | Conversation, files, and sessions | Active in the main codebase |
| `POST /api/chat` | Persistent conversational flow | Active; database-backed and file-aware |
| `POST /v1/study/rheumaai-reply` | Controlled evaluation of a de-identified note | Active; stateless, authenticated, and designed not to write RheumAI rows |
| `vanilla`, `rag`, `dag`, `quick-orvs`, `full-orvs` | Compare pipeline variants | Experimental; explicitly selected |
| Deep research | Longer evidence search and synthesis | Active as a separate flow |
| x402 | Optional payment for API clients | Optional, configuration-controlled integration |
## What happens during a clinical request
[Section titled “What happens during a clinical request”](#what-happens-during-a-clinical-request)
1. Access controls apply authentication, limits, and optional x402 enforcement.
2. RheumAI creates or retrieves the conversation and prepares request state.
3. Uploaded files are parsed before planning.
4. The planner selects tools and sources, such as PubMed, Semantic Scholar, local knowledge, the knowledge graph, or lab interpretation.
5. Retrieved sources are reranked for the question.
6. The response generator uses RheumAI’s clinical persona.
7. ORVS reviews the response when applicable. A failed score can trigger regeneration.
8. A separate gate checks retrieved PubMed IDs and removes unverified identifiers.
9. Ethical review can block a rejected output. If that review fails technically, current code passes the response through and records the failure.
10. The result returns to the clinician with the evidence available in that turn.
## What it does not do yet
[Section titled “What it does not do yet”](#what-it-does-not-do-yet)
* It does not independently produce a complete BIOBADAMEX record ready for submission.
* It does not score the **input note** against every required registry field.
* It does not return a structured distinction among missing, not applicable, and unknown data.
* It does not replace clinician review and approval.
The proposed BiobadamexAI integration uses RheumAI for reasoning and explanation. Registry schema, field-level provenance, completeness rules, and clinician approval remain in the controlled institutional workflow. See [Clinical notes and completeness](/en/rheumai/clinical-notes/) and the [current BiobadamexAI workflow](/en/biobadamexai/).
## Interpretation limits
[Section titled “Interpretation limits”](#interpretation-limits)
These pages describe code at commit `149508b3ab79da0b88c56deb63f6fb66e35d0e68`. A registered tool is not proof that every environment configures it or that it has clinical validation. Experimental paths and fallbacks are labeled throughout this section.
## Code evidence
[Section titled “Code evidence”](#code-evidence)
* [Server entry and routes](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/index.ts)
* [Main chat flow](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/routes/chat.ts)
* [Dynamic tool registry](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/tools/index.ts)
* [Stateless study route](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/study/rheumaai-reply-route.ts)
# Arquitectura
> Cómo los tres repos se conectan para producir un resultado del estudio — arms, request flow, scoring, persistencia.
> Esta página describe el estudio comparativo y sus arms. Para el producto clínico actual, consulta el [workflow de BiobadamexAI](/biobadamexai/).
Esta página no existe en ningún repo individual porque ninguno de los tres tiene el panorama completo — cada uno documenta su propio servicio, no cómo encaja con los otros dos. Este es ese mapa.
## Los tres arms, quién los corre
[Sección titulada «Los tres arms, quién los corre»](#los-tres-arms-quién-los-corre)
El código define tres “harnesses” (`STUDY_HARNESSES = ["plain", "rheumaai", "grader"]`, `pokta-care-monorepo/apps/api/src/services/biobadamex-study-harness.ts`). Un “arm” del estudio es en realidad `(modelo × harness × prompt opcional)` — no hay una lista fija de modelos hardcodeada, se configura en runtime vía la variable de entorno `STUDY_ARMS`.
| Harness | Cómo se invoca | Dónde vive el código |
| ---------- | ------------------------------------------------------------------------------ | ---------------------------------------------------------------------------- |
| `plain` | Llamada LLM **en proceso**, dentro del monorepo — sin repo externo involucrado | `pokta-care-monorepo` (`plainInvoker`) |
| `rheumaai` | HTTP hacia `RHEUMAI_STRUCTURED_URL` | servicio externo: `rheuma-ai-bioagent`, ruta `POST /v1/study/rheumaai-reply` |
| `grader` | HTTP hacia `GRADER_STRUCTURED_URL` | servicio externo: `pokta-grader-bioagent`, ruta `POST /v1/chat/completions` |
**El cuarto arm (LLM ajustado por otro reumatólogo) no es un harness separado.** Es el harness `plain` corriendo con un `promptId` distinto (`rheum-v1`), registrado en código pero con el prompt vacío/reservado — si se invoca, el código lanza un error en vez de correr silenciosamente con contenido incorrecto. No está “casi listo”, está sin construir todavía.
## Request flow
[Sección titulada «Request flow»](#request-flow)
```
flowchart TD
M["pokta-care-monorepo
(orquestador del estudio)"]
P["plain
(llamada LLM en proceso)"]
R["rheumaai
HTTP POST → rheuma-ai-bioagent
/v1/study/rheumaai-reply"]
G["grader
HTTP POST → pokta-grader-bioagent
/v1/chat/completions"]
O[("biobadamex_study_arm_outputs
una fila por nota × arm")]
S["scoring
(biobadamex-eval.ts)"]
A[("biobadamex_study_analysis
agregado, sin llave por nota")]
M --> P
M --> R
M --> G
P --> O
R --> O
G --> O
O --> S
S --> A
```
`rheumaai` y `grader` nunca se llaman entre sí ni comparten proceso — cada uno es un servicio HTTP independiente, desplegado por separado (Railway), con su propia autenticación por API key. El monorepo es el único que sabe que existen los tres; ninguno de los servicios sabe del otro (con una excepción cosmética: el comentario en `rheuma-ai-bioagent`’s `rheumaai-reply-route.ts` reconoce explícitamente que su técnica de auth fue copiada de `pokta-grader-bioagent`’s `structured.ts` — copiada, no importada, los repos no tienen dependencia en runtime).
## Sweep flow — de principio a fin
[Sección titulada «Sweep flow — de principio a fin»](#sweep-flow--de-principio-a-fin)
Secuencia real de `pokta-care-monorepo/apps/api/scripts/biobadamex-sweep-run.ts`:
1. **Cargar el corpus** — lista los `.docx` en `CORPUS_DIR` (ruta de filesystem, nunca dentro de un repo git). Deriva un `note_id = sha256(filename)[:12]`; el nombre real del archivo (que es PHI — nombre del paciente) nunca se loguea ni persiste.
2. **Resolver los arms** — lee `STUDY_ARMS` del entorno. Si está vacío, termina sin tocar nada.
3. **Modo dry-run opcional** — con `PLAN_ONLY=1`, imprime el plan (notas × arms) y termina — cero llamadas a proveedores, cero escritura a DB.
4. **Guard contra corridas previas sin drenar**, luego crea un `study_run`.
5. **Por cada nota**: la lee/parsea y arma un job por cada `(nota, arm)` — se encolan en `pg-boss`.
6. **Workers** toman los jobs y llaman al invoker correcto por arm, con un semáforo de concurrencia por proveedor.
7. **Poll cada 5s** hasta que todas las celdas estén “settled” (hay fila de salida, o se agotaron los reintentos), hasta 45 min.
8. **Reporte final**: cuenta esperado/settled/output/fallas, agrupado por causa de falla.
## Scoring — qué significa “completeness”
[Sección titulada «Scoring — qué significa “completeness”»](#scoring--qué-significa-completeness)
`compareField` (`biobadamex-eval.ts`) compara cada campo con un veredicto de cuatro estados, siempre usando un chequeo explícito de “desconocido” (nunca un chequeo falsy — así un `0` real nunca se confunde con “faltante”):
* **match** — ambos desconocidos, o los valores coinciden (para `patient.sex` específicamente, vía un catálogo tipo HL7 AdministrativeGender, no comparación exacta de string).
* **mismatch** — ambos conocidos, valores distintos.
* **missing** — el gold tiene valor, la extracción no.
* **unexpected** — la extracción tiene valor, el gold no.
**“Completeness” es una tasa de presencia, no una tasa de acierto.** Es `presente / total` sobre los campos requeridos que no son UNKNOWN — no factoriza si el valor es correcto contra el gold. La exactitud/acierto es una métrica separada (`variant_accuracy`, `delta_grader_minus_plain`). Esta distinción ya causó un bug real en el estudio (arm-2 estructurando la prosa del modelo en vez de la nota original inflaba completeness sin mejorar exactitud) — si vas a reportar un número, sé explícito sobre cuál de las dos métricas es.
## Dónde caen los resultados
[Sección titulada «Dónde caen los resultados»](#dónde-caen-los-resultados)
* **`biobadamex_study_arm_outputs`** — una fila por `(run_id, note_id, arm_id)`. Columnas relevantes: `model_id`, `harness`, `record` (el JSON extraído), `das28`, `missing_required`, `latency_ms`, tokens. Sin FK sobre `note_id` — a propósito.
* **`biobadamex_study_analysis`** — solo agregado, formato largo, **sin llave a nivel de nota** (a propósito, para prevenir re-identificación). Columnas: `run_id`, `arm_id`, `model_id`, `harness`, `scope`, `field_key`, `metric`, `numerator`, `n`, `value`, intervalos de confianza. Aquí es donde filtras por `scope="field"`, `metric="field_present_rate"` para ver el desglose campo por campo (ej. el campo `sex`).
## Corpus real de notas
[Sección titulada «Corpus real de notas»](#corpus-real-de-notas)
Vive fuera de todo repo, en una ruta de filesystem pasada por `CORPUS_DIR`. El script de censo (`biobadamex-instrument-census.ts`) verifica explícitamente que esa ruta **no** esté dentro de un working tree de git — “Real patient notes must never be committable” es un comentario literal en el código, no solo una convención. Ver [Reglas](/guardrails/) para las reglas de trabajo con datos sintéticos en su lugar.
# Arquitectura de BiobadamexAI
> Componentes, almacenes, colas y límites de confianza del flujo actual.
## Mapa del sistema
[Sección titulada «Mapa del sistema»](#mapa-del-sistema)
```
flowchart TB
UI["Consola React"] --> API["API Hono"]
M2M["Carga por servicio"] --> API
RAI["RheumAI audit proxy"] --> AUD["API de auditoría"]
AG["Agente clínico dedicado"] --> AGA["API de agente"]
API --> R2[("R2 · originales cifrados")]
API --> DB[("PostgreSQL · metadatos y borradores")]
API --> Q["pg-boss"]
AUD --> EXT["Extractor"]
AGA --> Q
Q --> EXT
EXT --> NEB["Nebius · nota desidentificada"]
EXT --> CORE["Esquema tri-state · DAS-28 · completitud"]
CORE --> DB
UI --> REVIEW["Revisión y corrección"]
REVIEW --> DB
DB --> SUB["Solicitud de envío"]
SUB --> Q2["Cola registry-submit"]
Q2 --> TENKI["MicroVM Tenki"]
TENKI --> REG["BIOBADAMEX ASP.NET"]
REG --> RECEIPT["idpac + resumen"]
RECEIPT --> DB
```
## Componentes
[Sección titulada «Componentes»](#componentes)
### Aplicación web
[Sección titulada «Aplicación web»](#aplicación-web)
La consola React concentra el flujo diario en una lista de notas. Los filtros muestran qué necesita atención y cada fila ofrece una acción contextual: extraer, revisar o ver. Las pantallas de revisión muestran el documento fuente, el registro estructurado, la corrección y la confirmación previa al envío.
### API
[Sección titulada «API»](#api)
La API Hono monta superficies separadas:
* `/api/biobadamex/intake`: texto pegado, archivos y carga por servicio.
* `/api/biobadamex/uploads`: inventario, extracción bajo demanda y recuperación del original.
* `/api/biobadamex/drafts`: cola de revisión, detalle, corrección, aprobación, rechazo y envío.
* `/api/biobadamex/audit`: calificación y envío para revisión desde RheumAI.
* `/api/biobadamex/agent`: acceso acotado para un agente que actúa por un médico.
* `/api/flujos`: línea de tiempo operativa con una proyección reducida.
Las rutas de consola exigen sesión, usuario resuelto y acceso de investigación. Las superficies máquina-a-máquina usan claves de servicio distintas.
### Almacenamiento
[Sección titulada «Almacenamiento»](#almacenamiento)
* R2 conserva el archivo original cifrado con AES-256-GCM.
* PostgreSQL guarda metadatos del archivo, trabajos, borradores, auditoría y resultados del envío.
* La copia de trabajo del registro preserva valores tri-state.
* `agentRecord` conserva la extracción original; `record` es la copia que el médico puede corregir.
* El texto fuente del borrador se almacena cifrado cuando existe una clave maestra válida.
### Cola
[Sección titulada «Cola»](#cola)
`pg-boss` separa extracción y registro. La API solo registra trabajadores cuando `BIOBADAMEX_QUEUE_ENABLED=1`. Extracción tiene reintentos y dead letter. Registro tiene otra cola y se procesa uno por trabajador, pero una instalación con varias instancias necesita coordinación adicional para garantizar serialización global.
### Paquete de dominio
[Sección titulada «Paquete de dominio»](#paquete-de-dominio)
`@pokta/biobadamex-registry` contiene el esquema, valores `unknown`, reglas por enfermedad, DAS-28, contratos de trabajos, mapa de controles externos y reglas de corrección y completitud.
### Trabajador de registro
[Sección titulada «Trabajador de registro»](#trabajador-de-registro)
El trabajador Playwright corre dentro de una microVM desechable. Inicia sesión, recorre `crdA` a `crdE`, obtiene el `idpac` después de guardar `crdA` y exige que el resumen final muestre ese identificador.
## Tres límites distintos
[Sección titulada «Tres límites distintos»](#tres-límites-distintos)
| Límite | Qué controla | Qué no demuestra |
| --------------------------- | ---------------------------------------- | -------------------------------------------- |
| Extracción | Convierte una nota en un registro tipado | Que todos los valores sean correctos |
| Revisión médica | Permite corregir y aprobar | Que la plataforma externa recibió cada campo |
| Confirmación del trabajador | Obtiene `idpac` y ve el resumen | Lectura campo por campo después del envío |
## Evidencia en código
[Sección titulada «Evidencia en código»](#evidencia-en-código)
* [Montaje de la API](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/api/src/index.ts)
* [Esquema del borrador](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/packages/db/src/schema/biobadamex-drafts.ts)
* [Configuración del pipeline](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/api/src/config/pipeline-config.ts)
* [Ensamblador de ciclo de vida](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/api/src/dal/biobadamex-lifecycle.ts)
* [Trabajador Playwright](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/biobadamex-registry-worker/src/driver.ts)
# Completitud y revisión médica
> Reglas por enfermedad, correcciones, aprobación, rechazo y confirmaciones clínicas.
## Dos barras que no deben mezclarse
[Sección titulada «Dos barras que no deben mezclarse»](#dos-barras-que-no-deben-mezclarse)
BiobadamexAI distingue:
1. **Campos que el registro externo exige:** confirmados por el comportamiento del formulario.
2. **Campos que el equipo clínico exige para una nota útil:** una barra clínica propia y dependiente de enfermedad.
La documentación y los análisis deben nombrar cada barra. Llamar “estándar del registro” a la segunda atribuiría a BIOBADAMEX una regla que no impone.
## Completitud por enfermedad
[Sección titulada «Completitud por enfermedad»](#completitud-por-enfermedad)
`requiredFieldsFor(record)` selecciona la variante antes de calcular faltantes:
* Artritis reumatoide: usa componentes clínicos confirmados para esa variante.
* Espondiloartritis: excluye TJC28, SJC28 y DAS-28; no fabrica BASDAI o ASDAS.
* Vasculitis: puede exigir BVAS según la regla confirmada.
* LES y esclerosis sistémica: conservan índices candidatos, pero no eligen uno mientras exista ambigüedad.
* Sjögren y variante desconocida mantienen una postura conservadora.
Algunos comentarios históricos en `extraction.ts` todavía describen una barra plana. El código ejecutable llama `requiredFieldsFor(record)` y esta documentación sigue la implementación probada.
## Borrador y procedencia
[Sección titulada «Borrador y procedencia»](#borrador-y-procedencia)
Cada borrador contiene `record` como copia corregible, `agentRecord` como snapshot inmutable, DAS-28 reconciliado, faltantes, fuente cifrada, modelo, huella del prompt, propietario, revisor y estado del envío.
La vista operativa reduce la diferencia agente-humano a conteos. No devuelve valores clínicos.
## Revisión
[Sección titulada «Revisión»](#revisión)
La pantalla permite ver la fuente, revisar filas específicas de enfermedad, corregir campos permitidos, recomputar DAS-28 y faltantes, revisar episodios biológicos y aprobar o rechazar.
Una corrección modifica `record`, nunca `agentRecord`, y genera auditoría. Cargar otro borrador o corregir reinicia confirmaciones.
## Qué bloquea cada paso
[Sección titulada «Qué bloquea cada paso»](#qué-bloquea-cada-paso)
| Paso | UI | Backend |
| -------- | --------------------------------------------- | ------------------------------------------------------------------------- |
| Aprobar | Exige confirmaciones, diagnóstico y episodios | Exige propiedad y estado `review_pending`; no recibe los ticks de UI |
| Corregir | Solo campos definidos | Valida path, tri-state, propiedad y estado; recomputa |
| Rechazar | Acción explícita | Exige propiedad y estado `review_pending` |
| Enviar | Drawer de confirmación | Exige propietario, `approved`, faltantes resueltos, flag y payload válido |
La checklist de aprobación vive principalmente en el cliente. El endpoint de aprobación no vuelve a validar esos ticks. El endpoint de envío sí vuelve a comprobar las condiciones que protegen la mutación externa.
La acción de UI dice “Aprobar y enviar al registro”, pero el endpoint solo cambia el estado a `approved`. El envío ocurre después desde la pantalla terminal. La copia no debe interpretarse como una mutación externa inmediata.
`No realizado` también es una exención cliente-side: desbloquea la aprobación, pero no cambia `unknown`, no limpia `missingRequired` y no deja auditoría. Por eso el envío posterior sigue bloqueado.
## Comorbilidades no documentadas
[Sección titulada «Comorbilidades no documentadas»](#comorbilidades-no-documentadas)
Cuando alguna permanece `unknown`, el drawer exige confirmar que lo no documentado debe registrarse como `No`. El backend exige la misma confirmación. Solo los `unknown` del bloque de comorbilidades se convierten a `false`; fechas, actividad, tratamientos y otros desconocidos no cambian. La normalización se persiste y audita.
El preview es completo para el **registro tipado interno**, no para todos los controles del formulario externo. Puede mostrar índices o episodios `requested/current` que el trabajador no envía; el trabajador solo coloca episodios `previous` en la cuadrícula de tratamientos previos.
## Interrupciones y aprendizaje
[Sección titulada «Interrupciones y aprendizaje»](#interrupciones-y-aprendizaje)
La calificación guarda `acted` o `dismissed`. Discrepar con la recomendación es un motivo de dismissal, no un tercer estado. La combinación borrador, prompt y actor se actualiza en vez de duplicarse.
Este ledger mide qué avisos valen una interrupción. No autoriza el envío ni sustituye la corrección clínica.
## Evidencia en código
[Sección titulada «Evidencia en código»](#evidencia-en-código)
* [Reglas por enfermedad](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/packages/biobadamex-registry/src/record.ts)
* [Completitud](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/packages/biobadamex-registry/src/completeness.ts)
* [Resultado de extracción](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/packages/biobadamex-registry/src/extraction.ts)
* [Rutas de borradores](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/api/src/routes/biobadamex-drafts.ts)
* [Pantalla de revisión](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/web/src/pages/BiobadamexReviewGate.tsx)
* [Drawer de envío](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/web/src/components/BiobadamexSubmitDrawer.tsx)
# Ingesta y extracción
> Cómo entran las notas, cómo se protegen y cómo se convierten en borradores.
## Formas de ingreso
[Sección titulada «Formas de ingreso»](#formas-de-ingreso)
| Entrada | Persistencia del original | Momento de extracción |
| --------------------------------------- | ------------------------------------------ | --------------------------------------------------------------------- |
| Texto pegado en consola | No crea archivo original | Inmediata, se encola al recibirlo |
| PDF, Word o texto en consola | Original cifrado en R2 + fila de metadatos | Bajo demanda desde el inventario |
| Carga por lotes con clave de servicio | Original cifrado por nota | Solo almacena; el médico extrae después |
| RheumAI “submit for review” | No crea fila de upload | Reextrae en servidor y crea un borrador retenido |
| Agente clínico dedicado | Usa solo uploads del médico configurado | Bajo demanda, con clave de agente |
| Documento del canal de estudio WhatsApp | Original cifrado en R2 | Inmediata cuando el canal, consentimiento y médico están configurados |
La respuesta de carga por lotes prueba almacenamiento, no extracción. Cada nota devuelve su propio resultado y una nota inválida no cancela sus hermanas.
## Archivos
[Sección titulada «Archivos»](#archivos)
1. El servidor valida formato, nombre y tamaño antes de decodificar por completo.
2. Extrae texto localmente. PDF usa `unpdf`, Word usa el parser de documentos y texto se decodifica directamente.
3. Cifra los bytes originales con AES-256-GCM.
4. Guarda el ciphertext en R2 y metadatos en PostgreSQL.
5. El inventario expone estado y nombre derivado, no el cuerpo de la nota.
6. Al solicitar extracción, la API comprueba alcance antes de obtener o descifrar el archivo.
El límite actual del archivo original es 20 MiB. Los nombres se limpian para impedir controles y separadores de ruta.
## Trabajo de extracción
[Sección titulada «Trabajo de extracción»](#trabajo-de-extracción)
El trabajador obtiene el texto, aplica desidentificación de mejor esfuerzo, ejecuta el modelo de `registro-extraction`, analiza JSON, convierte nulos a `unknown`, valida el esquema, calcula DAS-28, calcula faltantes por enfermedad y crea un borrador `review_pending` con fuente cifrada y atribución.
## Integridad
[Sección titulada «Integridad»](#integridad)
* El modelo no decide la semántica tri-state.
* Un cero real se conserva.
* `false` solo representa negación explícita, salvo comorbilidades confirmadas al enviar.
* DAS-28 reportado se compara; el valor autoritativo se calcula desde componentes cuando es posible.
* Episodios biológicos mantienen rol, fármaco, fechas y motivo separados.
* El solicitado sin fecha propia puede usar la fecha de solicitud o nota; no aplica a episodios actuales o previos.
## Privacidad y riesgo residual
[Sección titulada «Privacidad y riesgo residual»](#privacidad-y-riesgo-residual)
La desidentificación elimina identificadores genéricos y tokens del nombre cuando puede localizarlos. No es garantía de cumplimiento: un nombre libre sin etiqueta puede sobrevivir. Fechas clínicas, sexo, centro y médico pueden conservarse porque forman parte del registro.
El payload clínico de la cola de extracción puede contener texto en claro dentro de PostgreSQL. El camino de estudio cifra su nota antes de la cola, pero el camino clínico todavía no comparte esa protección.
## Reintentos y fallos
[Sección titulada «Reintentos y fallos»](#reintentos-y-fallos)
* Extracción permite cinco intentos con backoff.
* Al agotarlos, un dead-letter crea un fallo visible sin guardar el texto clínico.
* Una nota fallida o rechazada puede volver a extraerse.
* Un borrador pendiente o aprobado evita duplicados.
## Huecos operativos
[Sección titulada «Huecos operativos»](#huecos-operativos)
* Un fallo entre guardar el objeto R2 y crear la fila puede dejar un objeto cifrado huérfano.
* Un fallo de cola después de guardar el upload produce éxito parcial: la fuente existe aunque no haya trabajo.
* Si el borrador se crea y luego falla la auditoría, el reintento puede crear otro borrador porque la inserción no tiene llave de idempotencia por intento.
* Una nota pegada que termina en dead letter no tiene archivo original recuperable desde el inventario.
* Completar desde fuente llena solo escalares `unknown`; no fusiona arrays de episodios y no tiene UI clínica.
## Evidencia en código
[Sección titulada «Evidencia en código»](#evidencia-en-código)
* [Ingesta](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/api/src/routes/biobadamex-intake.ts)
* [Inventario y extracción](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/api/src/routes/biobadamex-uploads.ts)
* [Extractor](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/api/src/services/biobadamex-extract.ts)
* [Trabajos](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/packages/biobadamex-registry/src/jobs.ts)
* [Handlers de cola](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/api/src/services/biobadamex-jobs.ts)
# Integraciones y cambios seguros
> Autenticación, servicios externos, banderas y pruebas antes de modificar el workflow.
## Integraciones
[Sección titulada «Integraciones»](#integraciones)
| Integración | Uso | Límite |
| ------------------- | ---------------------------------------- | -------------------------------------------- |
| Privy | Sesión y atribución opcional | La clave de servicio autoriza rutas M2M |
| PostgreSQL/Drizzle | Metadatos, borradores, auditoría y colas | Contiene datos clínicos protegidos |
| R2 | Originales cifrados | Descifrado solo después de comprobar alcance |
| pg-boss | Extracción, dead letter y envío | Trabajadores solo si la cola está habilitada |
| Nebius | Extracción estructurada | Desidentificación no es garantía |
| RheumAI audit proxy | Calificar o enviar para revisión | No aprueba ni registra |
| Agente clínico | Inventario y extracción por médico | No aprueba ni envía |
| Tenki | MicroVM para Playwright | Requiere configuración y salida a internet |
| BIOBADAMEX | Formulario crdA–crdE | El mapa no cubre todos los controles |
## Configuración
[Sección titulada «Configuración»](#configuración)
Nombres únicamente, nunca valores:
* `BIOBADAMEX_QUEUE_ENABLED`
* `BIOBADAMEX_REGISTRY_SUBMIT_ENABLED`
* `NEBIUS_API_KEY`, `NEBIUS_BASE_URL`
* clave de ingesta M2M y `BIOBADAMEX_INGEST_MEDIC_ID`
* `BIOBADAMEX_AUDIT_API_KEY`
* clave del agente y médico `onBehalfOf`
* variables R2 y clave maestra de cifrado
* token, imagen y workspace de Tenki
* URL base y credenciales del registro
No copies valores a archivos, documentación, comandos, logs o PRs.
## Invariantes
[Sección titulada «Invariantes»](#invariantes)
1. Comprobar propietario antes de descifrar R2.
2. Mantener `agentRecord` inmutable.
3. Recomputar DAS-28 y faltantes tras corrección.
4. No convertir `unknown` a cero o `false` de forma general.
5. Separar aprobar de enviar.
6. Volver a comprobar propiedad, estado, faltantes, confirmación y flag al enviar.
7. Usar la versión del mapa del payload.
8. No equiparar `idpac` y resumen con read-back campo por campo.
9. Separar datos del estudio y workflow clínico.
10. Documentación no autoriza deploy, migración ni escritura externa.
## Mapa de cambio y pruebas
[Sección titulada «Mapa de cambio y pruebas»](#mapa-de-cambio-y-pruebas)
| Cambio | Archivos | Verificación mínima |
| --------------------- | -------------------------------- | ----------------------------------------- |
| Ingesta | intake/uploads | tests de intake, batch y uploads |
| Extracción | extract/config | extractor, fallos y esquema |
| Tri-state/completitud | paquete registry | suite completa del paquete |
| Revisión/corrección | drafts + UI | tests de drafts, preview y typecheck web |
| Cola | queue/jobs | jobs, dead-letter y lifecycle |
| Mapa/worker | field-map/resolver/filler/driver | suites registry y worker; bundle |
| Envío | submit + Tenki | mocks únicamente; nunca envío real |
| Auditoría/agentes | audit/agent | claves fail-closed, alcance y no mutación |
## Verificación observada
[Sección titulada «Verificación observada»](#verificación-observada)
```bash
pnpm --filter @pokta/biobadamex-registry test
pnpm --filter @pokta/biobadamex-registry-worker test
pnpm --filter @pokta/api test --
pnpm exec turbo run typecheck --force
```
En el commit documentado: registry 190 pruebas, worker 41, API enfocada 156 y typecheck forzado 12 tareas, todas pasaron. No se ejecutaron migraciones, despliegues, consultas clínicas ni escrituras al registro.
## Riesgos que deben permanecer visibles
[Sección titulada «Riesgos que deben permanecer visibles»](#riesgos-que-deben-permanecer-visibles)
* Texto clínico en claro dentro del payload de extracción.
* Desidentificación de mejor esfuerzo.
* Checklist de aprobación principalmente cliente-side.
* Worker best-effort para campos no fatales.
* Mapa incompleto y con ambigüedades.
* Sin read-back campo por campo.
* Registry-submit sin política específica de reintento.
* Descripciones médico-facing pendientes en pipeline config.
## Evidencia en código
[Sección titulada «Evidencia en código»](#evidencia-en-código)
* [Auditoría RheumAI](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/api/src/routes/biobadamex-audit.ts)
* [Agente clínico](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/api/src/routes/biobadamex-agent.ts)
* [Pipeline config](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/api/src/config/pipeline-config.ts)
* [Colas](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/api/src/lib/queue.ts)
* [Lifecycle](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/api/src/dal/biobadamex-lifecycle.ts)
# Envío al registro
> Condiciones, cola, sandbox, llenado de crdA–crdE y comprobación del resultado.
El envío es un proceso separado de la aprobación. No se dispara al aprobar y permanece deshabilitado salvo configuración explícita.
## Condiciones antes de encolar
[Sección titulada «Condiciones antes de encolar»](#condiciones-antes-de-encolar)
`POST /api/biobadamex/drafts/:id/submit` exige sesión, acceso de investigación, propiedad, estado `approved`, faltantes resueltos, `BIOBADAMEX_REGISTRY_SUBMIT_ENABLED=1`, confirmación de comorbilidades cuando aplica y payload válido con versión del mapa.
Un administrador puede leer borradores de otro médico con auditoría, pero no enviarlos. Un borrador sin propietario tampoco es enviable.
## Trabajo en cola
[Sección titulada «Trabajo en cola»](#trabajo-en-cola)
La API crea un trabajo con id del borrador, registro aprobado y versión del mapa. La cola usa lote de uno por trabajador, pero no configura reintentos específicos. Varias instancias necesitan coordinación adicional para garantizar una sola sesión externa global.
## Sandbox
[Sección titulada «Sandbox»](#sandbox)
Por cada trabajo, el servicio valida configuración, carga el bundle, crea una microVM Tenki, inyecta payload y credenciales mediante variables de proceso, ejecuta Playwright, analiza el resultado JSON y destruye la microVM al salir.
## Recorrido externo
[Sección titulada «Recorrido externo»](#recorrido-externo)
1. Login.
2. Llenar y guardar `crdA`.
3. Leer el `idpac` asignado.
4. Visitar y guardar `crdB`, `crdC`, `crdD` y `crdE`.
5. Abrir resumen.
6. Exigir el texto del mismo `idpac`.
`idpac` y resumen son fatales. Muchos campos individuales son best-effort: un control ausente o una opción no reconocida se registra y se omite. `succeeded` confirma paciente y resumen, no paridad campo por campo.
## Mapa externo
[Sección titulada «Mapa externo»](#mapa-externo)
El mapa está versionado. Cambiar selectores exige subir `REGISTRY_MAP_VERSION` y revalidar. Existen huecos conocidos: controles no modelados, radios no confirmados con valores distintos, ambigüedad BASDAI, marcas biológicas pendientes y campos de detalle/fecha no cubiertos.
No se debe describir como cobertura total del formulario.
## Tratamientos biológicos
[Sección titulada «Tratamientos biológicos»](#tratamientos-biológicos)
`crdE` acepta episodios con rol. Solo `previous` entra al bloque previo. El trabajador rechaza rol desconocido con fármaco, conflicto marca-sustancia, episodios duplicados que sobreescribirían fechas y fármacos sin mapeo confirmado.
## Resultado y errores
[Sección titulada «Resultado y errores»](#resultado-y-errores)
* Éxito: persiste `idpac`, `submittedAt`, estado `submitted` y auditoría.
* Rechazo del formulario: persiste `submitError`, audita y no repite el mismo envío.
* Fallo de infraestructura: lanza error; la cola actual no tiene política específica de reintento.
* La lectura posterior verifica resumen, no cada campo.
* Hay contratos de lectura acotada, pero no están montados en el flujo de submit de este commit.
## Riesgos de idempotencia y estado
[Sección titulada «Riesgos de idempotencia y estado»](#riesgos-de-idempotencia-y-estado)
* El endpoint descarta el id del trabajo y no usa singleton por borrador. Dos solicitudes antes del primer resultado pueden encolar dos mutaciones.
* Si `crdA` asigna `idpac` y falla una página posterior, el resultado fallido puede contenerlo en memoria, pero la persistencia guarda solo `submitError`. Un reintento puede volver a crear identidad.
* `summaryUrl` no se persiste.
* Un fallo de infraestructura antes de un resultado estructurado deja el borrador `approved` sin `submitError`; la UI no tiene un estado durable `queued/running`.
## Evidencia en código
[Sección titulada «Evidencia en código»](#evidencia-en-código)
* [Endpoint de envío](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/api/src/routes/biobadamex-drafts.ts)
* [Servicio Tenki](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/api/src/services/biobadamex-registry-submit.ts)
* [Handler de resultado](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/api/src/services/biobadamex-jobs.ts)
* [Mapa](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/packages/biobadamex-registry/src/field-map.ts)
* [Driver](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/biobadamex-registry-worker/src/driver.ts)
* [Llenado](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/biobadamex-registry-worker/src/field-filler.ts)
# Home
> Technical hub for the Grader study — architecture, onboarding, and links to every repo involved.
This site is the technical hub for people collaborating on the **comparative clinical-extraction study** and its core product, the **Grader**. If you just joined, start at [Getting Started](/en/getting-started/).
Restricted access: if you’re seeing this, you already have authorized access. The content here is as sensitive as the private repos it links to — treat it the same way, don’t forward it or publish it outside this group. See [Guardrails](/en/guardrails/) for details.
**Agents and LLMs:** start at [`/llms.txt`](/llms.txt). Use [`/llms-small.txt`](/llms-small.txt) for reduced context or [`/llms-full.txt`](/llms-full.txt) for the complete site.
## The project, in short
[Section titled “The project, in short”](#the-project-in-short)
We automate preparation of clinical records for a rheumatology biologics registry (the registry’s name and the collaborating entities stay in *stealth* mode — see [Guardrails](/en/guardrails/)). Physicians took their consultation note and had to re-enter it manually into the registry platform; the pay for doing so dropped from \~$30 to \~$12 USD per patient, and several physicians no longer do it. The bet: an engine that reads the note, prepares a structured draft, and helps the clinician correct and approve it before any controlled submission.
To back that bet with numbers we run a comparative study — several extraction “harnesses” over the same corpus of real notes:
1. **Plain LLM** — no specialized prompt, baseline. Lives in `pokta-care-monorepo`, not a separate service.
2. **RheumaAI** — an agent already trained in rheumatology. Own repo, own service.
3. **Grader** — the evolving product. Own repo, own service, forked from RheumaAI.
4. An optional fourth arm (an LLM tuned by another rheumatologist) — registered in code as a placeholder, not running yet.
How these pieces connect — who calls whom, where results land — is in [Architecture](/en/architecture/), not here; this page is just the “what’s what” map.
**North Star:** the end goal isn’t “read a Word doc and fill a form” — it’s a real-time copilot that, during the consultation, extracts the clinical information live and tells the physician “you forgot to ask X” before the visit ends. Retrospective note extraction is the training and validation for that engine, not the final product.
## The repos
[Section titled “The repos”](#the-repos)
| Repo | What it is | Service |
| ----------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------ | ------------------------------- |
| [`pokta-grader-bioagent`](https://github.com/poktalabs/pokta-grader-bioagent) | The Grader — arm 3 of the study | Railway, `grader.poktacare.com` |
| [`rheuma-ai-bioagent`](https://github.com/poktalabs/rheuma-ai-bioagent) | RheumaAI — arm 2 of the study | Railway, `rheumai.xyz` |
| [`pokta-care-monorepo`](https://github.com/poktalabs/pokta-care-monorepo) | Study orchestrator: arm 1, scoring, persistence, and the product app (`app.poktacare.com`) | Vercel + Railway |
Full detail, with direct links to each one’s README/CONTRIBUTING, in [Repos & Guides](/en/repos/).
## Team
[Section titled “Team”](#team)
* **Dr. Erick Zamora Tehozol** — Rheumatology and clinical validation
* **Ing. Ángel Meléndez Córdoba** — Engineering
## Where to go next
[Section titled “Where to go next”](#where-to-go-next)
* **[Getting Started](/en/getting-started/)** — access checklist, where to start depending on which repo you’ll work in.
* **[BiobadamexAI](/en/biobadamexai/)** — current workflow: intake, extraction, clinical review, and controlled submission.
* **[Architecture](/en/architecture/)** — how the three repos connect: who calls whom, how it’s measured, where results land.
* **[Repos & Guides](/en/repos/)** — link directory.
* **[Guardrails](/en/guardrails/)** — the non-negotiable rules. Read them before your first change.
# Architecture
> How the three repos connect to produce a study result — arms, request flow, scoring, persistence.
> This page describes the comparative study and its arms. For the current clinical product, see the [BiobadamexAI workflow](/en/biobadamexai/).
This page doesn’t exist in any single repo because none of the three has the full picture — each documents its own service, not how it fits with the other two. This is that map.
## The three arms, who runs them
[Section titled “The three arms, who runs them”](#the-three-arms-who-runs-them)
The code defines three “harnesses” (`STUDY_HARNESSES = ["plain", "rheumaai", "grader"]`, `pokta-care-monorepo/apps/api/src/services/biobadamex-study-harness.ts`). A study “arm” is really `(model × harness × optional prompt)` — there’s no fixed hardcoded model list, it’s configured at runtime via the `STUDY_ARMS` environment variable.
| Harness | How it’s invoked | Where the code lives |
| ---------- | ------------------------------------------------------------------------ | ----------------------------------------------------------------------------- |
| `plain` | LLM call **in-process**, inside the monorepo — no external repo involved | `pokta-care-monorepo` (`plainInvoker`) |
| `rheumaai` | HTTP to `RHEUMAI_STRUCTURED_URL` | external service: `rheuma-ai-bioagent`, route `POST /v1/study/rheumaai-reply` |
| `grader` | HTTP to `GRADER_STRUCTURED_URL` | external service: `pokta-grader-bioagent`, route `POST /v1/chat/completions` |
**The fourth arm (LLM tuned by another rheumatologist) isn’t a separate harness.** It’s the `plain` harness running with a different `promptId` (`rheum-v1`), registered in code but with the prompt empty/reserved — if invoked, the code throws instead of silently running with wrong content. It’s not “almost ready,” it’s not built yet.
## Request flow
[Section titled “Request flow”](#request-flow)
```
flowchart TD
M["pokta-care-monorepo
(study orchestrator)"]
P["plain
(in-process LLM call)"]
R["rheumaai
HTTP POST → rheuma-ai-bioagent
/v1/study/rheumaai-reply"]
G["grader
HTTP POST → pokta-grader-bioagent
/v1/chat/completions"]
O[("biobadamex_study_arm_outputs
one row per note × arm")]
S["scoring
(biobadamex-eval.ts)"]
A[("biobadamex_study_analysis
aggregated, no per-note key")]
M --> P
M --> R
M --> G
P --> O
R --> O
G --> O
O --> S
S --> A
```
`rheumaai` and `grader` never call each other or share a process — each is an independent HTTP service, deployed separately (Railway), with its own API-key auth. The monorepo is the only piece that knows all three exist; neither service knows about the other (with one cosmetic exception: a comment in `rheuma-ai-bioagent`’s `rheumaai-reply-route.ts` explicitly acknowledges its auth technique was copied from `pokta-grader-bioagent`’s `structured.ts` — copied, not imported; the repos have no runtime dependency).
## Sweep flow — start to finish
[Section titled “Sweep flow — start to finish”](#sweep-flow--start-to-finish)
Actual sequence from `pokta-care-monorepo/apps/api/scripts/biobadamex-sweep-run.ts`:
1. **Load the corpus** — lists the `.docx` files in `CORPUS_DIR` (a filesystem path, never inside a git repo). Derives a `note_id = sha256(filename)[:12]`; the real filename (which is PHI — the patient’s name) is never logged or persisted.
2. **Resolve the arms** — reads `STUDY_ARMS` from the environment. If empty, exits without touching anything.
3. **Optional dry-run mode** — with `PLAN_ONLY=1`, prints the plan (notes × arms) and exits — zero provider calls, zero DB writes.
4. **Guards against undrained previous runs**, then creates a `study_run`.
5. **For each note**: reads/parses it and builds a job per `(note, arm)` — enqueued in `pg-boss`.
6. **Workers** pick up jobs and call the right invoker for each arm, with a per-provider concurrency semaphore.
7. **Polls every 5s** until every cell is “settled” (there’s an output row, or retries are exhausted), up to 45 min.
8. **Final report**: counts expected/settled/output/failures, grouped by failure cause.
## Scoring — what “completeness” means
[Section titled “Scoring — what “completeness” means”](#scoring--what-completeness-means)
`compareField` (`biobadamex-eval.ts`) compares each field with a four-state verdict, always using an explicit “unknown” check (never a falsy check — so a real `0` is never confused with “missing”):
* **match** — both unknown, or the values match (for `patient.sex` specifically, via an HL7 AdministrativeGender-style catalog, not exact string comparison).
* **mismatch** — both known, values differ.
* **missing** — the gold has a value, the extraction doesn’t.
* **unexpected** — the extraction has a value, the gold doesn’t.
**“Completeness” is a presence rate, not an accuracy rate.** It’s `present / total` over the required fields that aren’t UNKNOWN — it doesn’t factor in whether the value is correct against the gold. Accuracy is a separate metric (`variant_accuracy`, `delta_grader_minus_plain`). This distinction already caused a real bug in the study (arm-2 structuring the model’s prose instead of the original note inflated completeness without improving accuracy) — if you’re reporting a number, be explicit about which of the two metrics it is.
## Where results land
[Section titled “Where results land”](#where-results-land)
* **`biobadamex_study_arm_outputs`** — one row per `(run_id, note_id, arm_id)`. Relevant columns: `model_id`, `harness`, `record` (the extracted JSON), `das28`, `missing_required`, `latency_ms`, tokens. No FK on `note_id` — on purpose.
* **`biobadamex_study_analysis`** — aggregated only, long format, **no note-level key** (on purpose, to prevent re-identification). Columns: `run_id`, `arm_id`, `model_id`, `harness`, `scope`, `field_key`, `metric`, `numerator`, `n`, `value`, confidence intervals. This is where you filter by `scope="field"`, `metric="field_present_rate"` to see the field-by-field breakdown (e.g. the `sex` field).
## Real note corpus
[Section titled “Real note corpus”](#real-note-corpus)
Lives outside any repo, at a filesystem path passed via `CORPUS_DIR`. The census script (`biobadamex-instrument-census.ts`) explicitly verifies that path is **not** inside a git working tree — “Real patient notes must never be committable” is a literal comment in the code, not just a convention. See [Guardrails](/en/guardrails/) for the rules on working with synthetic data instead.
# BiobadamexAI architecture
> Components, storage, queues, and trust boundaries in the current workflow.
## System map
[Section titled “System map”](#system-map)
```
flowchart TB
UI["React console"] --> API["Hono API"]
M2M["Service intake"] --> API
RAI["RheumAI audit proxy"] --> AUD["Audit API"]
AG["Dedicated clinician agent"] --> AGA["Agent API"]
API --> R2[("R2 · encrypted originals")]
API --> DB[("PostgreSQL · metadata and drafts")]
API --> Q["pg-boss"]
AUD --> EXT["Extractor"]
AGA --> Q
Q --> EXT
EXT --> NEB["Nebius · de-identified note"]
EXT --> CORE["Tri-state schema · DAS-28 · completeness"]
CORE --> DB
UI --> REVIEW["Review and correction"]
REVIEW --> DB
DB --> SUB["Submission request"]
SUB --> Q2["registry-submit queue"]
Q2 --> TENKI["Tenki microVM"]
TENKI --> REG["BIOBADAMEX ASP.NET"]
REG --> RECEIPT["idpac + summary"]
RECEIPT --> DB
```
## Components
[Section titled “Components”](#components)
### Web application
[Section titled “Web application”](#web-application)
The React console uses one note list for daily work. Filters show what needs attention. Review screens show source, structured record, corrections, treatment episodes, and a separate submission confirmation.
### API
[Section titled “API”](#api)
Mounted surfaces include intake, uploads, drafts, RheumAI audit, dedicated agent, and PHI-light lifecycle routes. Console routes require authentication, resolved user, and research access. Machine-to-machine routes use separate service keys.
### Storage
[Section titled “Storage”](#storage)
R2 holds encrypted original files. PostgreSQL stores metadata, jobs, drafts, audit, and submission outcomes. `agentRecord` is the immutable extraction; `record` is the clinician-correctable copy. Draft source text is encrypted when master-key configuration is available.
### Queue
[Section titled “Queue”](#queue)
`pg-boss` separates extraction from registration. Workers register only when `BIOBADAMEX_QUEUE_ENABLED=1`. Extraction has retries and dead letter. Registration is a separate queue processed one at a time per worker. Multiple app instances need extra coordination for global serialization.
### Domain package
[Section titled “Domain package”](#domain-package)
`@pokta/biobadamex-registry` owns the record schema, `unknown` semantics, disease rules, DAS-28, job contracts, external field map, corrections, and completeness.
### Registry worker
[Section titled “Registry worker”](#registry-worker)
A Playwright worker runs in a disposable microVM. It logs in, traverses `crdA` through `crdE`, reads `idpac` after `crdA`, and requires the final summary to show the same id.
## Three distinct boundaries
[Section titled “Three distinct boundaries”](#three-distinct-boundaries)
| Boundary | Controls | Does not prove |
| ------------------- | ---------------------------- | -------------------------------------- |
| Extraction | Turns note into typed record | Every value is correct |
| Clinical review | Corrects and approves | External platform received every field |
| Worker confirmation | Reads idpac and summary | Field-by-field read-back |
## Code evidence
[Section titled “Code evidence”](#code-evidence)
* [API mounting](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/api/src/index.ts)
* [Draft schema](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/packages/db/src/schema/biobadamex-drafts.ts)
* [Pipeline config](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/api/src/config/pipeline-config.ts)
* [Lifecycle assembler](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/api/src/dal/biobadamex-lifecycle.ts)
* [Playwright worker](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/biobadamex-registry-worker/src/driver.ts)
# Completeness and clinical review
> Disease-specific rules, corrections, approval, rejection, and clinical confirmations.
## Two bars that must remain separate
[Section titled “Two bars that must remain separate”](#two-bars-that-must-remain-separate)
BiobadamexAI distinguishes registry-enforced fields from the team’s disease-specific clinical documentation bar. Every report must name the bar. Calling the clinical bar a registry standard would attribute an internal rule to BIOBADAMEX.
## Disease-specific completeness
[Section titled “Disease-specific completeness”](#disease-specific-completeness)
`requiredFieldsFor(record)` selects the variant before calculating missing fields:
* Rheumatoid arthritis uses confirmed components for that variant.
* Spondyloarthritis excludes TJC28, SJC28, and DAS-28 and does not invent BASDAI or ASDAS.
* Vasculitis may require BVAS under the confirmed rule.
* SLE and systemic sclerosis retain candidate instruments without choosing one while clinical ambiguity remains.
* Sjögren and unknown variant remain conservative.
Some historical comments in `extraction.ts` still describe a flat bar. Executable code calls `requiredFieldsFor(record)`, and this documentation follows the tested implementation.
## Draft and provenance
[Section titled “Draft and provenance”](#draft-and-provenance)
Each draft has a correctable `record`, immutable `agentRecord`, reconciled DAS-28, missing fields, encrypted source, model, prompt fingerprint, ownership, reviewer, and submission state. The operational view reduces agent-human differences to counts and does not return clinical values.
## Review
[Section titled “Review”](#review)
The clinician can view source, review disease-specific rows, correct allowed fields, recompute DAS-28 and missing fields, review biologic episodes, then approve or reject. Corrections modify `record`, never `agentRecord`, and create audit entries.
## What blocks each step
[Section titled “What blocks each step”](#what-blocks-each-step)
| Step | UI | Backend |
| ------- | ----------------------------------------------- | ------------------------------------------------------------------------------- |
| Approve | Requires confirmations, diagnosis, and episodes | Requires ownership and `review_pending`; does not receive UI ticks |
| Correct | Defined fields only | Validates path, tri-state, ownership, and state; recomputes |
| Reject | Explicit action | Requires ownership and `review_pending` |
| Submit | Confirmation drawer | Requires owner, `approved`, no missing required fields, flag, and valid payload |
The approval checklist is mainly client-side. The approval endpoint does not revalidate its ticks. The submission endpoint rechecks conditions that protect the external mutation.
The UI action says “Approve and send to registry,” but the endpoint only changes state to `approved`. Submission happens later from the terminal-state screen. The copy must not be read as an immediate external mutation.
`Not performed` is also a client-side waiver: it unlocks approval but does not change `unknown`, clear `missingRequired`, or create audit evidence. Later submission therefore remains blocked.
## Undocumented comorbidities
[Section titled “Undocumented comorbidities”](#undocumented-comorbidities)
If any remain `unknown`, the drawer and backend require explicit confirmation before converting only those comorbidities to `false`. Dates, activity, treatments, and other unknowns stay unchanged. The normalization is persisted and audited.
The preview is complete for the **internal typed record**, not every external form control. It may show indices or `requested/current` episodes the worker does not submit; the worker sends only `previous` episodes to the prior-treatment grid.
## Interruption learning
[Section titled “Interruption learning”](#interruption-learning)
Feedback records `acted` or `dismissed`. Disagreement is a dismissal reason, not a third state. Draft, prompt, and actor are upserted rather than duplicated. This ledger measures useful interruptions; it does not authorize submission.
## Code evidence
[Section titled “Code evidence”](#code-evidence)
* [Disease rules](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/packages/biobadamex-registry/src/record.ts)
* [Completeness](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/packages/biobadamex-registry/src/completeness.ts)
* [Extraction result](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/packages/biobadamex-registry/src/extraction.ts)
* [Draft routes](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/api/src/routes/biobadamex-drafts.ts)
* [Review screen](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/web/src/pages/BiobadamexReviewGate.tsx)
* [Submit drawer](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/web/src/components/BiobadamexSubmitDrawer.tsx)
# Intake and extraction
> How notes enter, how they are protected, and how they become drafts.
## Intake modes
[Section titled “Intake modes”](#intake-modes)
| Input | Original persistence | Extraction timing |
| ------------------------------- | ----------------------------------- | ------------------------------------------------------------- |
| Pasted console text | No original-file row | Immediate queueing |
| PDF, Word, or text file | Encrypted R2 object + metadata | On demand from inventory |
| Service-key batch | Encrypted per-note original | Storage only; clinician extracts later |
| RheumAI submit for review | No upload row | Server re-extracts and creates held draft |
| Dedicated clinician agent | Configured clinician’s uploads only | On demand with agent key |
| WhatsApp study-channel document | Encrypted original in R2 | Immediate when channel, consent, and clinician are configured |
Batch response proves storage, not extraction. One invalid note does not abort its siblings.
## Files
[Section titled “Files”](#files)
The server validates format, filename, and size, extracts text locally, encrypts original bytes with AES-256-GCM, stores ciphertext in R2, and stores metadata in PostgreSQL. Scope is checked before R2 retrieval or decryption. Current file ceiling is 20 MiB.
## Extraction job
[Section titled “Extraction job”](#extraction-job)
The worker obtains note text, applies best-effort de-identification, calls the configured `registro-extraction` model, parses JSON, converts nulls to `unknown`, validates schema, calculates DAS-28, calculates disease-matched missing fields, and creates a `review_pending` draft with encrypted source and attribution.
## Integrity
[Section titled “Integrity”](#integrity)
* Model does not define tri-state semantics.
* Real zero remains zero.
* `false` means explicit negation except clinician-confirmed comorbidity normalization at submission.
* Reported DAS-28 is compared; computation uses components where possible.
* Biologic episodes retain role, drug, dates, and stop reason.
* Requested biologic without its own date may use request/note date; current and previous episodes do not.
## Privacy and residual risk
[Section titled “Privacy and residual risk”](#privacy-and-residual-risk)
De-identification removes generic identifiers and patient-name tokens when found. It is not a compliance guarantee. Unlabelled names can survive, while clinically required dates and provider context may remain.
Clinical extraction job payloads can contain plaintext note text in PostgreSQL. Study jobs encrypt their note before queueing, but the clinical path does not yet share that protection.
## Retries and failures
[Section titled “Retries and failures”](#retries-and-failures)
Extraction permits five attempts with backoff. Exhaustion creates an app-visible dead letter without note text. Failed or rejected notes may be re-extracted; pending or approved drafts prevent duplicate extraction.
## Operational gaps
[Section titled “Operational gaps”](#operational-gaps)
* Failure between R2 storage and metadata insert can leave an orphan encrypted object.
* Queue failure after upload persistence produces partial success: source exists without a job.
* Draft creation followed by audit failure can create duplicate drafts on retry because success insertion has no attempt idempotency key.
* A pasted note that dead-letters has no recoverable original in inventory.
* Source completion fills unknown scalar fields only; it does not merge episode arrays and has no clinician UI.
## Code evidence
[Section titled “Code evidence”](#code-evidence)
* [Intake](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/api/src/routes/biobadamex-intake.ts)
* [Inventory and extraction](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/api/src/routes/biobadamex-uploads.ts)
* [Extractor](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/api/src/services/biobadamex-extract.ts)
* [Jobs](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/packages/biobadamex-registry/src/jobs.ts)
* [Queue handlers](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/api/src/services/biobadamex-jobs.ts)
# Integrations and safe changes
> Authentication, external services, flags, and checks before changing the workflow.
## Integrations
[Section titled “Integrations”](#integrations)
| Integration | Use | Boundary |
| ------------------- | ------------------------------------------ | -------------------------------------------- |
| Privy | Session and optional attribution | Service key authorizes M2M routes |
| PostgreSQL/Drizzle | Metadata, drafts, audit, queues | Protected clinical data |
| R2 | Encrypted originals | Decrypt only after scope check |
| pg-boss | Extraction, dead letter, submission | Workers only when enabled |
| Nebius | Structured extraction | De-identification is not a guarantee |
| RheumAI audit proxy | Grade or submit for review | Cannot approve or register |
| Clinician agent | Inventory and extraction for one clinician | Cannot approve or submit |
| Tenki | Playwright microVM | Requires explicit config and outbound access |
| BIOBADAMEX | crdA–crdE form | Map does not cover every control |
## Configuration names
[Section titled “Configuration names”](#configuration-names)
Names only, never values:
* `BIOBADAMEX_QUEUE_ENABLED`
* `BIOBADAMEX_REGISTRY_SUBMIT_ENABLED`
* `NEBIUS_API_KEY`, `NEBIUS_BASE_URL`
* M2M intake key and optional `BIOBADAMEX_INGEST_MEDIC_ID`
* `BIOBADAMEX_AUDIT_API_KEY`
* clinician-agent key and configured `onBehalfOf` clinician
* R2 variables and encryption master key
* Tenki token, image, and workspace
* registry base URL and credentials
Never copy values into files, docs, commands, logs, or PRs.
## Invariants
[Section titled “Invariants”](#invariants)
1. Check owner before R2 decryption.
2. Keep `agentRecord` immutable.
3. Recompute DAS-28 and missing fields after correction.
4. Never convert `unknown` to zero or `false` generally.
5. Keep approval separate from submission.
6. Recheck ownership, state, missing fields, confirmation, and flag at submission.
7. Use payload map version.
8. Do not equate idpac/summary with field-level read-back.
9. Keep study data separate from clinical workflow.
10. Documentation never authorizes deploy, migration, or external write.
## Change and test map
[Section titled “Change and test map”](#change-and-test-map)
| Change | Files | Minimum verification |
| ---------------------- | -------------------------- | --------------------------------------------- |
| Intake | intake/uploads | intake, batch, upload tests |
| Extraction | extract/config | extractor, failure, schema tests |
| Tri-state/completeness | registry package | full registry suite |
| Review/correction | drafts + UI | draft, preview tests, web typecheck |
| Queue | queue/jobs | jobs, dead letter, lifecycle |
| Map/worker | map/resolver/filler/driver | registry and worker suites; bundle |
| Submission | submit + Tenki | mocks only; never live submit |
| Audit/agents | audit/agent | fail-closed keys, scope, no external mutation |
## Observed validation
[Section titled “Observed validation”](#observed-validation)
```bash
pnpm --filter @pokta/biobadamex-registry test
pnpm --filter @pokta/biobadamex-registry-worker test
pnpm --filter @pokta/api test --
pnpm exec turbo run typecheck --force
```
At the documented commit: registry 190 tests, worker 41, focused API 156, and forced typecheck 12 tasks all passed. No migrations, deployments, clinical queries, or registry writes ran.
## Risks that must remain visible
[Section titled “Risks that must remain visible”](#risks-that-must-remain-visible)
* Plaintext clinical text in extraction queue payload.
* Best-effort de-identification.
* Approval checklist mainly client-side.
* Best-effort nonfatal field filling.
* Incomplete and ambiguous external map.
* No field-by-field read-back.
* No registry-submit-specific retry policy.
* Pending clinician-facing descriptions in pipeline config.
## Code evidence
[Section titled “Code evidence”](#code-evidence)
* [RheumAI audit](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/api/src/routes/biobadamex-audit.ts)
* [Clinician agent](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/api/src/routes/biobadamex-agent.ts)
* [Pipeline config](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/api/src/config/pipeline-config.ts)
* [Queues](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/api/src/lib/queue.ts)
* [Lifecycle](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/api/src/dal/biobadamex-lifecycle.ts)
# Registry submission
> Preconditions, queue, sandbox, crdA–crdE filling, and outcome confirmation.
Submission is separate from approval. Approval never auto-submits, and submission stays disabled without explicit configuration.
## Preconditions
[Section titled “Preconditions”](#preconditions)
`POST /api/biobadamex/drafts/:id/submit` requires session, research access, ownership, `approved` state, resolved required fields, `BIOBADAMEX_REGISTRY_SUBMIT_ENABLED=1`, comorbidity confirmation when needed, and a valid map-versioned payload.
Admins may audit-read another clinician’s draft but cannot submit it. Unowned drafts are not submittable.
## Queue
[Section titled “Queue”](#queue)
The API enqueues draft id, approved record, and map version. Registry submit uses batch size one per worker but no specific retry policy. Multiple app instances need extra coordination to guarantee one global external session.
## Sandbox
[Section titled “Sandbox”](#sandbox)
For each job, the service validates configuration, loads the worker bundle, creates a disposable Tenki microVM, injects payload and credentials through process environment, runs Playwright, parses structured JSON, and disposes the microVM.
## External flow
[Section titled “External flow”](#external-flow)
1. Log in.
2. Fill and save `crdA`.
3. Read assigned `idpac`.
4. Visit and save `crdB`, `crdC`, `crdD`, and `crdE`.
5. Open summary.
6. Require text for the same `idpac`.
`idpac` and summary are fatal checks. Many individual fields are best effort: missing controls or unmatched options are logged and skipped. `succeeded` confirms patient and summary, not field-by-field parity.
## External map
[Section titled “External map”](#external-map)
The control map is versioned. Selector changes require a version bump and revalidation. Known gaps include unmapped controls, radio ambiguities, BASDAI ambiguity, biologic brands awaiting validation, and uncovered detail/date fields. Do not describe it as total form coverage.
## Biologic treatments
[Section titled “Biologic treatments”](#biologic-treatments)
`crdE` accepts role-aware episodes. Only `previous` enters prior treatments. The worker rejects unknown role with identified drug, brand-substance conflict, duplicate control overwrite, and unmapped drugs.
## Outcomes
[Section titled “Outcomes”](#outcomes)
* Success persists `idpac`, `submittedAt`, `submitted` state, and audit.
* Form rejection persists `submitError`, audits, and does not repeat the same submission.
* Infrastructure failure throws, but the queue has no specific registry-submit retry policy.
* Current read-back confirms summary, not each field.
* Bounded read contracts exist but are not mounted in this commit’s submit flow.
## Idempotency and state risks
[Section titled “Idempotency and state risks”](#idempotency-and-state-risks)
* The endpoint drops the job id and uses no per-draft singleton. Two requests before the first outcome can enqueue two external mutations.
* If `crdA` assigns an `idpac` and a later page fails, the failed result may carry it in memory, but persistence stores only `submitError`. Retry can repeat identity creation.
* `summaryUrl` is not persisted.
* Infrastructure failure before structured result leaves the draft `approved` without `submitError`; the UI has no durable `queued/running` state.
## Code evidence
[Section titled “Code evidence”](#code-evidence)
* [Submit endpoint](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/api/src/routes/biobadamex-drafts.ts)
* [Tenki service](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/api/src/services/biobadamex-registry-submit.ts)
* [Outcome handler](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/api/src/services/biobadamex-jobs.ts)
* [Field map](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/packages/biobadamex-registry/src/field-map.ts)
* [Driver](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/biobadamex-registry-worker/src/driver.ts)
* [Field filler](https://github.com/poktalabs/pokta-care-monorepo/blob/b852f056d527d5bcec72b67256e5af98d386b323/apps/biobadamex-registry-worker/src/field-filler.ts)
# Getting Started
> Access checklist and where to start depending on which repo you'll work in.
## Access — checklist
[Section titled “Access — checklist”](#access--checklist)
* [ ] Access to the repo you’ll work in (see table below).
* [ ] Access to `poktalabs/pokta-care-monorepo` — the benchmark scripts and scoring live there, even if your main work is in another repo.
* [ ] Signed up in the app (`app.poktacare.com`) — signup is open, you just need to log in with Gmail, no allowlist for basic access. Access to the research surfaces (the study dashboard and the BIOBADAMEX console) is gated by a separate email allowlist — if your work needs it, ask whoever onboarded you.
* [ ] Access to the “grader persona” document in Google Drive (relevant if you’re working on the Grader’s extraction/retrieval).
* [ ] Access to the chat group with Dr. Erick Zamora, to coordinate technical handoffs with the person who made most of the rheumatology modifications.
* [ ] Access to this site (`docs.poktacare.com`) — you already have it if you’re reading this.
If any of these is still pending, say so directly to whoever onboarded you — this isn’t something you should work around.
## Where to start
[Section titled “Where to start”](#where-to-start)
**You’re working on the Grader (extraction, prompts, retrieval, arm 3):** Start at [`pokta-grader-bioagent/CONTRIBUTING.md`](https://github.com/poktalabs/pokta-grader-bioagent/blob/main/CONTRIBUTING.md). That file has repo-specific context, your first concrete task, and environment setup (`dev/SETUP.md`). Come back to [Architecture](/en/architecture/) here to understand how your work fits into the full study.
**You’re working on RheumaAI (arm 2):** Start at the `README.md` of `rheuma-ai-bioagent`. The Grader repo is a fork of this one — if you already know one, the other will feel familiar (same Bun runtime, same LLM library, same tools architecture).
**You’re working on the study harness, scoring, or the product app:** Start at `pokta-care-monorepo`, specifically `apps/api/scripts/biobadamex-*.ts` for the benchmark and `apps/api/src/services/biobadamex-eval.ts` for scoring. Read [Architecture](/en/architecture/) first — it has the full map of how this repo calls the other two.
**Not sure / your work crosses several repos:** Read the full [Architecture](/en/architecture/) page before touching code. It’s the only page that explains the system end to end.
## The honest starting point
[Section titled “The honest starting point”](#the-honest-starting-point)
The Grader and RheumaAI are forks of a framework called BioAgents (Bio/Dezy team) that no longer has active upstream maintenance. The infrastructure was set up and handed to Dr. Erick Zamora, who built the rheumatology modifications (full homomorphic encryption, specialized articles) on top of it. **No one on the team fully understands 100% of this code today** — part of your real job, in either repo, is reverse-engineering inherited pieces. That’s normal, not a sign of poor prior documentation.
You’re not training or fine-tuning a model — everything here today is orchestration of already-trained LLMs with prompts, extraction schemas, and (optionally) context retrieval via embeddings.
# Non-negotiable rules
> The rules that don't get negotiated — stealth, PHI, study integrity, push/merge. Canonical copy; individual repos link here.
This is the **canonical copy** of these rules — each repo’s `CONTRIBUTING.md` links here instead of repeating the full text. If you’re citing or updating them, do it on this page.
These rules aren’t bureaucracy: they’re what keeps a well-intentioned change from causing a real problem (legal, clinical-trust, or study-validity). Read them before your first line of code, not after.
## 1. Stealth / treat as confidential
[Section titled “1. Stealth / treat as confidential”](#1-stealth--treat-as-confidential)
The registry and the collaborating physicians/entities are in *stealth* mode. There’s no signed NDA, but the treatment is as if there were: **no public surface may name the registry, the collaborating physicians, or the entities involved, or imply their endorsement/participation.** This includes public repos, posts, third-party demos, portfolio, LinkedIn.
This site is access-gated, not public — but the content is just as sensitive as the private repos it links to. Don’t forward it, don’t take screenshots to share outside the authorized group, don’t assume “gated” means “casual.” If you’re unsure whether something counts as a public surface, ask before publishing.
## 2. Real patient data never enters a repo
[Section titled “2. Real patient data never enters a repo”](#2-real-patient-data-never-enters-a-repo)
The corpus of 110 notes is real patient information (the filenames are patient names) and stays outside git on purpose — not even in a private repo, not “just for testing.” Work with synthetic notes, never with the real corpus.
Example synthetic notes already exist: `src/routes/persona-isolation.test.ts` and `persona-runtime-guard.test.ts` in `pokta-grader-bioagent` (one note at a time), and the hardcoded examples in the monorepo’s `biobadamex-sweep-dry-run.ts`. None are at the scale of 110 — if you need more synthetic variety, generate invented notes, never ask for the real ones for local development.
## 3. Study integrity — the easiest rule to break unintentionally
[Section titled “3. Study integrity — the easiest rule to break unintentionally”](#3-study-integrity--the-easiest-rule-to-break-unintentionally)
**Any improvement whose number will back the study is measured on held-out data, never on the same 110 notes the study reports results against.**
Why it matters: if you tune a change while watching how well it does against those 110 notes, and then report that same number as a study result, you’re training against the test set. The number goes up because you specifically tuned for that set — not because the system extracts better in general — and that invalidates the comparison between arms, which is the entire point of the exercise.
This is easy to violate without bad intent: “I improved the score” feels good and is the closest metric at hand. The rule exists precisely because it’s the most natural trap to fall into.
Improving any of the three systems as a product is the goal, and that’s perfectly fine. The nuance is where you measure the number you’re going to show as a study result.
## 4. Push and merge
[Section titled “4. Push and merge”](#4-push-and-merge)
The three code repos (`pokta-grader-bioagent`, `rheuma-ai-bioagent`, `pokta-care-monorepo`) prohibit working directly on `main` or pushing directly — branches + PR + human review, always. This documentation site follows the same rule. Third-party forks (like BioAgents) are never pushed back to upstream/third-party remotes.
## What’s still undecided
[Section titled “What’s still undecided”](#whats-still-undecided)
* Compensation / formal collaboration terms for external contributors.
* Coordinating pending handoffs with Dr. Erick Zamora.
* Whether the optional fourth study arm ever gets built.
If any of these is blocking you, say so directly to whoever onboarded you — don’t improvise a solution for something that isn’t yours to decide.
# Repos & Guides
> Link directory — each repo, its README/CONTRIBUTING, and where it's deployed.
## `pokta-grader-bioagent`
[Section titled “pokta-grader-bioagent”](#pokta-grader-bioagent)
The Grader — arm 3 of the study, the evolving product.
* Repo: [github.com/poktalabs/pokta-grader-bioagent](https://github.com/poktalabs/pokta-grader-bioagent)
* Onboarding: [`CONTRIBUTING.md`](https://github.com/poktalabs/pokta-grader-bioagent/blob/main/CONTRIBUTING.md)
* Environment setup: [`dev/SETUP.md`](https://github.com/poktalabs/pokta-grader-bioagent/blob/main/dev/SETUP.md), [`dev/getting-started.md`](https://github.com/poktalabs/pokta-grader-bioagent/blob/main/dev/getting-started.md) (generic to the BioAgents framework)
* Engineering conventions: [`AGENTS.md`](https://github.com/poktalabs/pokta-grader-bioagent/blob/main/AGENTS.md)
* Deployed: Railway, `grader.poktacare.com`
* Runtime: Bun + Elysia
## `rheuma-ai-bioagent`
[Section titled “rheuma-ai-bioagent”](#rheuma-ai-bioagent)
RheumaAI — arm 2 of the study, an agent already trained in rheumatology. Fork origin of the Grader.
* Repo: [`poktalabs/rheuma-ai-bioagent`](https://github.com/poktalabs/rheuma-ai-bioagent) (private)
* Onboarding: `README.md` (explains routes, tools, state model, LLM library)
* Deployed: Railway, `rheumai.xyz`
* Runtime: Bun + Elysia — same stack as the Grader
## `pokta-care-monorepo`
[Section titled “pokta-care-monorepo”](#pokta-care-monorepo)
Study orchestrator (arm 1 in-process, calls the other two arms over HTTP, scoring, persistence) and the actual product app.
* Repo: [`poktalabs/pokta-care-monorepo`](https://github.com/poktalabs/pokta-care-monorepo)
* Benchmark scripts: `apps/api/scripts/biobadamex-*.ts`
* Scoring: `apps/api/src/services/biobadamex-eval.ts`
* Arm invocation: `apps/api/src/services/biobadamex-study-arms.ts` and `biobadamex-study-harness.ts`
* Deployed: Vercel (`app.poktacare.com` / `api.poktacare.com`) + Railway
* Stack: TypeScript, Turbo, Drizzle, React
## This site
[Section titled “This site”](#this-site)
* Repo: [`poktalabs/pokta-care-docs`](https://github.com/poktalabs/pokta-care-docs)
* Stack: Astro + Starlight, deployed on Cloudflare Pages under `docs.poktacare.com`
* Access: gated — see [Guardrails](/en/guardrails/)
# RheumAI architecture
> Components, data flow, and integration points in the current system.
## System map
[Section titled “System map”](#system-map)
```
flowchart TB
UI["Preact web app"] --> API["Elysia API"]
EXT["External client"] --> API
STUDY["BiobadamexAI study harness"] --> SR["Stateless study route"]
API --> ACCESS["Auth · rate limits · optional x402"]
ACCESS --> SETUP["Conversation and state"]
SETUP --> PLAN["PLANNING"]
PLAN --> TOOLS["Clinical tools and sources"]
TOOLS --> RET["DAG / RAG / graph"]
RET --> REPLY["REPLY or HYPOTHESIS"]
REPLY --> ORVS["ORVS + PMID control"]
ORVS --> ETH["Ethical review"]
ETH --> UI
SR --> PLAN2["PLANNING with tool cap"]
PLAN2 --> TOOLS2["Allowed tools"]
TOOLS2 --> REPLY2["REPLY"]
REPLY2 --> STUDY
SETUP --> DB[("Supabase")]
TOOLS --> OBJ[("S3-compatible storage")]
TOOLS --> LIT["PubMed · Semantic Scholar · OpenScholar · local knowledge"]
```
## Components
[Section titled “Components”](#components)
### Web client
[Section titled “Web client”](#web-client)
The Preact client manages sessions, authentication, uploads, chat requests, citations, and payment states. The backend serves the bundle and uses a single-page-app fallback.
### Elysia API
[Section titled “Elysia API”](#elysia-api)
`src/index.ts` mounts authentication, configuration, chat, deep research, community, search, cryptography, and study routes. It also serves the client and health endpoints.
### State and persistence
[Section titled “State and persistence”](#state-and-persistence)
Normal chat creates users, conversations, messages, and request state in Supabase. Files can be written to S3-compatible storage, while parsed content enters request state. `/api/chat` is therefore not stateless.
The study route uses nil UUIDs and omits message and state identifiers. This gates off writes by design. It accepts de-identified text, can pin one provider/model, and returns attribution, tool use, latency, and token usage.
### Planner and tools
[Section titled “Planner and tools”](#planner-and-tools)
The tool registry discovers folders under `src/tools/` and loads enabled `index.ts` modules. `TOOL__ENABLED` variables can override registration. The planner chooses sources and the primary action.
Relevant families include:
* Evidence: PubMed, Semantic Scholar, OpenScholar, local knowledge, and knowledge graph.
* Clinical: clinical reasoning, lab interpretation, insurance summary, and protocols.
* Synthesis: `REPLY`, `HYPOTHESIS`, reflection, and ethical review.
* Infrastructure: files, data analysis, storage, and medical x402 services.
A folder’s presence does not prove that it loads. Registration logs failures and continues with available tools.
### Retrieval and ranking
[Section titled “Retrieval and ranking”](#retrieval-and-ranking)
The pipeline can combine embeddings, local graph retrieval, and semantic search. `DAG-RETRIEVAL` reranks documents for the question. Some modes fall back to a direct model response when useful evidence is unavailable.
### Models
[Section titled “Models”](#models)
The LLM library exposes a common interface across providers. `resilientChatCompletion` implements fallbacks. A study context can pin all observable model calls to one provider/model and disable cross-model fallback.
## Two flows that must not be confused
[Section titled “Two flows that must not be confused”](#two-flows-that-must-not-be-confused)
| Aspect | Normal chat | Study route |
| ------------------ | ------------------------------------------------------- | ------------------------------------------ |
| Endpoint | `/api/chat` | `/v1/study/rheumaai-reply` |
| Persistence | Yes | No, by design |
| Files | Yes | Note text only |
| Session | Persistent conversation | Isolated request |
| Auth | Session, client, or x402 depending on config | Dedicated bearer; closed when key is unset |
| Tools | Planner-selected | Allowlist and cost ceiling |
| Post-response ORVS | Yes in normal flow; explicit comparison modes available | No explicit final ORVS call in this route |
| Output | Conversational response | Response, attribution, and turn telemetry |
## External dependencies
[Section titled “External dependencies”](#external-dependencies)
* Supabase for conversations, messages, state, and search.
* S3-compatible storage for files when configured.
* Configurable LLM providers.
* PubMed, Semantic Scholar, OpenScholar, and local corpora for evidence.
* Optional medical x402 services.
* Railway and `rheumai.xyz` according to repository configuration.
## Code evidence
[Section titled “Code evidence”](#code-evidence)
* [Server and route mounting](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/index.ts)
* [Chat route](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/routes/chat.ts)
* [Tool orchestration](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/services/chat/tools.ts)
* [Comparison pipeline](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/services/chat/pipeline.ts)
* [Tool registry](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/tools/index.ts)
* [LLM abstraction](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/llm/provider.ts)
* [Study route](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/study/rheumaai-reply-route.ts)
# Clinical notes and completeness
> What RheumAI can do today and how it should connect to the controlled BiobadamexAI flow.
> This page defines the RheumAI boundary. Current intake, completeness, review, and submission behavior lives in the [BiobadamexAI workflow](/en/biobadamexai/).
## The problem
[Section titled “The problem”](#the-problem)
A note can be well written and still omit a required data point. It can also contain a correct fact that software fails to structure. The system must separate three questions:
1. What does the note say?
2. What does the clinical workflow or registry require?
3. What should the clinician ask or correct before approval?
## Current RheumAI capability
[Section titled “Current RheumAI capability”](#current-rheumai-capability)
RheumAI can:
* receive clinical text and files in chat;
* parse PDF, spreadsheet, CSV, JSON, Markdown, and text files;
* use vision for some PDFs and images when configured;
* classify clinical question types and detect some alert terms;
* retrieve evidence and generate clinical interpretation;
* ask for missing context through persona and prompt instructions;
* review response quality with ORVS;
* process one de-identified note through a stateless study route.
These capabilities can help reveal gaps, but their main output today is narrative text.
## What RheumAI does not expose as a contract
[Section titled “What RheumAI does not expose as a contract”](#what-rheumai-does-not-expose-as-a-contract)
The repository does not yet return a structured payload containing:
* every BIOBADAMEX-required field;
* the extracted candidate value;
* the source text span;
* `present`, `missing`, `unknown`, or `not_applicable` state;
* the clinical rule that makes the field applicable;
* a suggested question for the clinician;
* clinician approval or rejection.
The study route returns content, provider, model, tools, tokens, latency, and fallback status. It does not return a note-completeness matrix. It also has no deterministic de-identification check, note-specific length limit, final ORVS pass, ethical review, or clinician approval. The caller must satisfy those duties before and after the route.
## Recommended integration design
[Section titled “Recommended integration design”](#recommended-integration-design)
The institutional flow should keep responsibilities separate:
| Stage | Owner | Output |
| -------------------- | ----------------------------------- | --------------------------------------------------- |
| Intake | BiobadamexAI | Controlled note and internal identifier |
| Extraction | Registry structured service | Candidate fields with provenance |
| Completeness | Diagnosis-specific BIOBADAMEX rules | Present, missing, unknown, or not-applicable fields |
| Clinical explanation | RheumAI | Prioritized gaps and useful clinician questions |
| Response review | ORVS and deterministic controls | Accuracy, safety, citation, and coverage flags |
| Human review | Authorized clinician | Correction and explicit approval |
| Submission | Controlled registry service | Approved data only, with read-back |
## Central rule
[Section titled “Central rule”](#central-rule)
RheumAI should **detect and ask, not invent**. If the note lacks a fact, the integration preserves missing or unknown status according to the clinical rule. Model inference never becomes registry data automatically.
## Completeness flow
[Section titled “Completeness flow”](#completeness-flow)
1. The extractor identifies candidate values and preserves supporting text.
2. The rule engine selects fields applicable to the diagnosis.
3. The system computes completeness as presence across applicable fields. Accuracy and completeness remain separate metrics.
4. RheumAI turns gaps into clear, prioritized questions.
5. ORVS reviews the **RheumAI response**, not the chart.
6. The clinician confirms, corrects, or marks data unavailable.
7. Only the approved version can reach the registry service.
## Safeguards
[Section titled “Safeguards”](#safeguards)
* Use synthetic or de-identified notes outside the authorized clinical environment.
* Remember that normal chat can persist messages and files. Long PDFs that resemble articles can also enter the local knowledge directory under current rules.
* Never place real notes in repositories, docs, logs, or fixtures.
* Preserve field-level provenance.
* Do not treat valid `0`, `false`, or dates as missing through generic truthiness checks.
* Keep missing, unknown, and not applicable distinct where clinical rules require it.
* Do not let ORVS or narrative output bypass clinician review.
* Record the model and rules version that created each draft.
## Readiness criteria
[Section titled “Readiness criteria”](#readiness-criteria)
Do not call the integration ready until it has:
* a versioned input/output contract;
* diagnosis-specific completeness rules under clinical review;
* synthetic tests for present and missing fields;
* clinical review of suggested questions;
* an audit that the route does not unexpectedly write or log PHI;
* service-to-service authentication;
* a fallback that preserves the draft without submission when RheumAI fails.
## Code evidence
[Section titled “Code evidence”](#code-evidence)
* [Stateless study route](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/study/rheumaai-reply-route.ts)
* [Study route tests](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/study/rheumaai-reply-route.test.ts)
* [File parsing](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/tools/file-upload/index.ts)
* [Lab interpretation](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/tools/lab-interpretation/index.ts)
* [Clinical reasoning](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/tools/clinical-reasoning/index.ts)
* [ORVS](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/tools/verification/index.ts)
# Integrations and safe changes
> External contracts, configuration, and minimum checks before modifying RheumAI.
## Main integrations
[Section titled “Main integrations”](#main-integrations)
| Integration | Use | Expected failure behavior |
| ------------------------------ | ------------------------------------------------- | --------------------------------------------------------------------- |
| Supabase | Users, conversations, messages, state, and search | Persistent chat can fail or degrade |
| S3-compatible storage | Conversation files | Text may parse while persistent upload is skipped |
| LLM providers | Planning, tools, response, and evaluation | `resilientChatCompletion` can fall back outside pinned study requests |
| PubMed | Biomedical evidence and verified PMIDs | Output should declare that no PMID was verified |
| Semantic Scholar / OpenScholar | Additional literature | Tool failure is logged and the pipeline may continue |
| Local knowledge / graph | Specialized retrieval | Another retrieval method or direct response may be used |
| BiobadamexAI study harness | Calls stateless route with de-identified note | 401 without valid bearer; 503 when key is unset |
| x402 | Optional external API payment | 402 or 429 depending on payment and limits |
## Behavior-defining variables
[Section titled “Behavior-defining variables”](#behavior-defining-variables)
This list names configuration, never values:
* Server: `PORT`, `HOST`, `CORS_ORIGINS`.
* Study: `RHEUMAI_STUDY_API_KEY`, `CORS_STUDY_ORIGINS`, `RHEUMAI_STUDY_TOOL_ALLOWLIST`, `RHEUMAI_STUDY_MAX_PAID_CALLS`.
* Models: `REPLY_LLM_PROVIDER`, `REPLY_LLM_MODEL`, `HYP_LLM_PROVIDER`, `HYP_LLM_MODEL`, `STRUCTURED_LLM_MODEL`.
* ORVS: `ORVS_ENABLED`, `ORVS_VERSION`.
* Storage and database: provider-specific variables; never copy values into docs.
* x402: `X402_ENABLED`, environment, payment address, and facilitator credentials.
## Invariants
[Section titled “Invariants”](#invariants)
1. The study route stays stateless and closed when its key is missing.
2. Note bodies never enter logs or error responses.
3. A study pin forces one provider/model and disables cross-model fallback.
4. Cost-bearing tools are excluded by default and the default paid ceiling is zero.
5. Published PMIDs come from the current turn’s retrieved set.
6. Normal clinical chat resolves to `REPLY` unless an explicit protocol or research flow applies.
7. A RheumAI failure never authorizes registry submission.
8. Only a clinician authorizes final clinical data.
## Change and test map
[Section titled “Change and test map”](#change-and-test-map)
| Change | Review | Minimum checks |
| ----------------------- | ----------------------------------------------------------- | ---------------------------------------------------------------------------- |
| Route mounting or CORS | `src/index.ts` and Elysia plugins | `bun run build`, non-mutating API smoke |
| Chat order or fallbacks | `src/routes/chat.ts`, `src/services/chat/*` | tool tests and synthetic smoke |
| Study route | `src/study/*` | `bun test src/study/rheumaai-reply-route.test.ts src/study/tool-cap.test.ts` |
| Providers/models | `src/llm/*` | `bun test src/llm/resilient.test.ts` and pin tests |
| ORVS or citations | `src/tools/verification/*` and PMID gate | unit tests, allowed and fabricated PMID cases |
| Files or labs | `src/tools/file-upload/*`, `src/tools/lab-interpretation/*` | lab tests with synthetic fixtures |
| Dynamic tools | `src/tools/index.ts` | start registry and inspect loaded/failed tools |
| Client | `client/src/*` | `bun run check`, `bun run build` |
## Recommended local verification
[Section titled “Recommended local verification”](#recommended-local-verification)
```bash
bun install --frozen-lockfile
bun run check
bun run build
bun test src/study/rheumaai-reply-route.test.ts \
src/study/tool-cap.test.ts \
src/tools/lab-interpretation/lab-interpretation.test.ts \
src/tools/knowledgeGraph/knowledgeGraph.test.ts
```
At the documented commit, the focused set above passed **26 tests with 0 failures** after a frozen install. Tool registration also emitted warnings for unconfigured tools and a folder without `index.ts`; focused tests are not a complete server-health check.
Smoke tests that call external services require valid configuration. Routine changes must not run them against production or with real notes.
## Code states
[Section titled “Code states”](#code-states)
* **Active:** imported by `src/index.ts`, chat, or a loaded tool and supported by observed tests.
* **Configurable:** active only when a variable or provider is present.
* **Experimental:** comparison modes, benchmarks, and non-default scripts.
* **Placeholder or broken:** registered folder without an `index.ts`, failed import, or integration requiring missing credentials.
* **Legacy:** retained for compatibility but not preferred.
Do not document a folder as working just because it exists. Confirm import, configuration, and tests.
Automated coverage does not include the main chat, deep research, authentication, community, Telegram, or x402 routes, and there are no client tests. Changes in those areas need synthetic smoke checks in addition to the existing tests.
## Code evidence
[Section titled “Code evidence”](#code-evidence)
* [Routes and CORS](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/index.ts)
* [Chat and PMID controls](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/routes/chat.ts)
* [Tool registry](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/tools/index.ts)
* [Model fallbacks](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/llm/resilient.ts)
* [Study route](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/study/rheumaai-reply-route.ts)
* [Study tool cap](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/study/tool-cap.ts)
# ORVS in RheumAI
> What ORVS checks, how current code uses it, and its limits.
ORVS means **Optimistic Response Verification System**. Erick Zamora published the reference document as a preprint on March 28, 2026, with DOI `10.55277/researchhub.qq2fdvw9`.\[1]
In RheumAI, ORVS is a **post-generation review**. It evaluates the model’s response. It does not directly validate whether the input clinical note contains every field BIOBADAMEX requires.
## Preprint and current code
[Section titled “Preprint and current code”](#preprint-and-current-code)
They are not the same contract. The preprint describes a four-dimension rubric, up to three cycles, and human escalation.\[1] Normal chat uses the legacy six-dimension version, attempts one regeneration, and has no ORVS human-escalation queue. V2 modes are experimental extensions: Quick uses four dimensions and Deep uses eight. Quick’s weights resemble the preprint, but current `TMP` and `RSC` meanings changed to temporal reasoning and therapeutic completeness.
## Conceptual flow
[Section titled “Conceptual flow”](#conceptual-flow)
1. RheumAI generates a response from the question and retrieved evidence.
2. An evaluator model scores several dimensions.
3. A low score can trigger regeneration or a targeted addendum.
4. A separate control compares cited PubMed IDs against IDs retrieved in that turn.
5. The system returns the result or applies a fallback when time runs out.
## Implementations in the repository
[Section titled “Implementations in the repository”](#implementations-in-the-repository)
### Normal chat
[Section titled “Normal chat”](#normal-chat)
Normal chat calls v1 verification for clinical requests and non-trivial questions. It scores six dimensions from 0–100:
* citation accuracy;
* clinical accuracy;
* specificity;
* evidence alignment;
* response completeness;
* absence of contradictions.
The threshold is 70. A failed score causes one regeneration attempt with evaluator flags. The regenerated response is not sent through ORVS again in that normal path.
### Explicit comparison modes
[Section titled “Explicit comparison modes”](#explicit-comparison-modes)
A client that supplies `pipelineMode` can select `quick-orvs` or `full-orvs`:
| Mode | Dimensions | Purpose |
| ------------- | ----------------------------------------------------------------------- | ------------ |
| Quick ORVS v2 | clinical accuracy, safety, temporal reasoning, therapeutic completeness | Fast review |
| Deep ORVS v2 | the four above plus completeness, citations, clarity, and bias | Wider review |
V2 computes a 0–100 composite and can regenerate or append targeted corrections up to two times.
## Citation controls
[Section titled “Citation controls”](#citation-controls)
The chat flow adds deterministic PubMed rules:
* Retrieved PubMed evidence without an inline PMID can trigger regeneration.
* Only PMIDs retrieved during the current turn are permitted.
* A remaining disallowed identifier is replaced with `[unverified PMID]`.
This is more deterministic than asking a model whether a citation looks plausible.
## Important limits
[Section titled “Important limits”](#important-limits)
* ORVS evaluates generated output. It does not prove that a recommendation is clinically correct.
* The evaluator is also a model and can fail.
* The preprint describes a proposal and retrospective experiments using constructed scenarios.\[1] It does not demonstrate clinical outcomes, regulatory approval, or safety for autonomous use.
* Current parse or API errors can return a result marked `passed` with error flags. This is fail-open behavior, not successful validation.
* `ORVS_ENABLED=false` disables verification.
* The stateless study route does not explicitly run final ORVS after `REPLY`.
* ORVS “completeness” asks whether the **response covers the question**. It is not chart or registry completeness.
## Relationship to note completeness
[Section titled “Relationship to note completeness”](#relationship-to-note-completeness)
BiobadamexAI needs two separate evaluations:
1. **Note completeness:** applicable registry fields are present, missing, unknown, or not applicable.
2. **Response quality:** RheumAI’s explanation is accurate, safe, evidence-aware, clear, and useful.
ORVS can support the second. The first requires an explicit clinical schema and field-level provenance. See [Clinical notes and completeness](/en/rheumai/clinical-notes/).
## Code evidence
[Section titled “Code evidence”](#code-evidence)
* [ORVS implementation](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/tools/verification/index.ts)
* [Evaluation prompts](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/tools/verification/prompts.ts)
* [Chat use and PMID controls](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/routes/chat.ts)
* [Comparison modes](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/services/chat/pipeline.ts)
## Sources
[Section titled “Sources”](#sources)
\[1] — Optimistic Response Verification System > “Optimistic Response Verification System” > “March 28, 2026”
# Primeros pasos
> Checklist de acceso y por dónde empezar según en qué repo vas a trabajar.
## Acceso — checklist
[Sección titulada «Acceso — checklist»](#acceso--checklist)
* [ ] Acceso al repo en el que vas a trabajar (ver tabla abajo).
* [ ] Acceso a `poktalabs/pokta-care-monorepo` — ahí viven los scripts del benchmark y el scoring, aunque tu trabajo principal sea en otro repo.
* [ ] Alta en la app (`app.poktacare.com`) — el registro está abierto, solo necesitas entrar con tu Gmail, no hay lista blanca para el acceso básico. El acceso a las superficies de investigación (dashboard del estudio y consola BIOBADAMEX) sí se controla por una lista blanca de correos aparte — si tu trabajo la requiere, pídesela a quien te dio de alta.
* [ ] Acceso al documento “grader persona” en Google Drive (relevante si trabajas en extracción/retrieval del Grader).
* [ ] Acceso al grupo de chat con el Dr. Erick Zamora, para coordinar handoffs técnicos con quien hizo la mayoría de las modificaciones de reumatología.
* [ ] Acceso a este sitio (`docs.poktacare.com`) — ya lo tienes si estás leyendo esto.
Si alguno sigue pendiente, dilo directamente a quien te dio de alta — no es algo que debas resolver rodeando el acceso.
## Por dónde empezar
[Sección titulada «Por dónde empezar»](#por-dónde-empezar)
**Vas a trabajar en el Grader (extracción, prompts, retrieval, arm 3):** Empieza en [`pokta-grader-bioagent/CONTRIBUTING.md`](https://github.com/poktalabs/pokta-grader-bioagent/blob/main/CONTRIBUTING.md). Ese archivo tiene el contexto específico del repo, tu primera misión concreta y el setup de entorno (`dev/SETUP.md`). Vuelve a [Arquitectura](/architecture/) aquí para entender cómo tu trabajo encaja en el estudio completo.
**Vas a trabajar en RheumaAI (arm 2):** Empieza en el `README.md` de `rheuma-ai-bioagent`. El repo del Grader es un fork de este — si ya conoces uno, el otro te va a resultar familiar (mismo runtime Bun, misma librería LLM, misma arquitectura de tools).
**Vas a trabajar en el harness del estudio, scoring, o la app de producto:** Empieza en `pokta-care-monorepo`, específicamente `apps/api/scripts/biobadamex-*.ts` para el benchmark y `apps/api/src/services/biobadamex-eval.ts` para el scoring. Lee [Arquitectura](/architecture/) primero — ahí está el mapa completo de cómo este repo llama a los otros dos.
**No estás seguro / tu trabajo cruza varios repos:** Lee [Arquitectura](/architecture/) completo antes de tocar código. Es la única página que explica el sistema de punta a punta.
## El punto de partida honesto
[Sección titulada «El punto de partida honesto»](#el-punto-de-partida-honesto)
El Grader y RheumaAI son forks de un framework llamado BioAgents (equipo Bio/Dezy) que ya no tiene mantenimiento activo upstream. La infraestructura se montó y se le entregó al Dr. Erick Zamora, quien hizo las modificaciones de reumatología (encriptación homomórfica completa, artículos especializados) encima. **Nadie en el equipo entiende hoy el 100% de este código internamente** — parte de tu trabajo real, en cualquiera de los dos repos, es ingeniería inversa de piezas heredadas. Es normal, no una señal de mala documentación previa.
No estás entrenando ni afinando un modelo — todo lo que hay hoy es orquestación de LLMs ya entrenados con prompts, schemas de extracción y (opcionalmente) recuperación de contexto vía embeddings.
# Reglas no negociables
> Las reglas que no se negocian — stealth, PHI, integridad del estudio, push/merge. Copia canónica; los repos individuales enlazan aquí.
Esta es la **copia canónica** de estas reglas — los `CONTRIBUTING.md` de cada repo enlazan aquí en vez de repetir el texto completo. Si vas a citarlas o actualizarlas, hazlo en esta página.
Estas reglas no son burocracia: son las que evitan que un cambio bien intencionado cause un problema real (legal, de confianza médica, o de validez del estudio). Léelas antes de tu primera línea de código, no después.
## 1. Stealth / trato como confidencial
[Sección titulada «1. Stealth / trato como confidencial»](#1-stealth--trato-como-confidencial)
El registro y los médicos/entidades colaboradoras están en modo *stealth*. No hay NDA firmado, pero el trato es como si lo hubiera: **ninguna superficie pública puede nombrar el registro, a los médicos colaboradores, ni a las entidades involucradas, ni insinuar su respaldo/participación.** Esto incluye repos públicos, posts, demos a terceros, portafolio, LinkedIn.
Este sitio está gateado por acceso, no es público — pero el contenido es igual de sensible que los repos privados que enlaza. No lo reenvíes, no tomes screenshots para compartir fuera del grupo autorizado, no asumas que “gateado” significa “casual”. Si tienes duda sobre si algo cuenta como superficie pública, pregunta antes de publicar.
## 2. Los datos de pacientes reales nunca entran a un repo
[Sección titulada «2. Los datos de pacientes reales nunca entran a un repo»](#2-los-datos-de-pacientes-reales-nunca-entran-a-un-repo)
El corpus de 110 notas es información real de pacientes (los nombres de archivo son nombres de pacientes) y está fuera de git a propósito — ni siquiera en un repo privado, ni “solo para pruebas”. Trabaja con notas sintéticas, nunca con el corpus real.
Ya existen notas sintéticas de ejemplo: `src/routes/persona-isolation.test.ts` y `persona-runtime-guard.test.ts` en `pokta-grader-bioagent` (una nota a la vez), y los ejemplos hardcodeados en `biobadamex-sweep-dry-run.ts` del monorepo. Ninguno es a escala de 110 — si necesitas más variedad sintética, genera notas inventadas, nunca pidas las reales para desarrollo local.
## 3. Integridad del estudio — la más fácil de romper sin querer
[Sección titulada «3. Integridad del estudio — la más fácil de romper sin querer»](#3-integridad-del-estudio--la-más-fácil-de-romper-sin-querer)
**Cualquier mejora cuyo número vaya a respaldar el estudio se mide en datos held-out, nunca en las mismas 110 notas contra las que el estudio reporta resultados.**
Por qué importa: si ajustas un cambio mirando qué tan bien le va contra esas 110 notas, y luego reportas ese mismo número como resultado del estudio, estás entrenando contra el set de prueba. El número sube porque ajustaste específicamente para ese conjunto — no porque el sistema extraiga mejor en general — y eso invalida la comparación entre arms, que es todo el punto del ejercicio.
Esto es fácil de violar sin mala intención: “mejoré el score” se siente bien y es la métrica más a la mano. La regla existe justamente porque es la trampa más natural en la que caer.
Mejorar cualquiera de los tres sistemas como producto es el objetivo, y está perfecto. El matiz es dónde mides el número que vas a mostrar como resultado del estudio.
## 4. Push y merge
[Sección titulada «4. Push y merge»](#4-push-y-merge)
Los tres repos de código (`pokta-grader-bioagent`, `rheuma-ai-bioagent`, `pokta-care-monorepo`) prohíben trabajar directo en `main` o hacer push directo — ramas + PR + revisión humana, siempre. Este sitio de documentación sigue la misma regla. Los forks de terceros (como BioAgents) nunca se empujan de vuelta a remotos upstream/terceros.
## Lo que todavía no está decidido
[Sección titulada «Lo que todavía no está decidido»](#lo-que-todavía-no-está-decidido)
* Compensación / términos formales de colaboración para contribuidores externos.
* Coordinación de handoffs pendientes con el Dr. Erick Zamora.
* Si el cuarto arm opcional del estudio se llega a construir.
Si alguno de estos te bloquea, dilo directamente a quien te haya dado de alta — no improvises una solución para un punto que no es tuyo.
# Flujo del MVP
> El recorrido de principio a fin del clínico en RheumAI — pantalla por pantalla, con cada compuerta, punto de decisión y estado de construcción (enviado, tras bandera, o sin construir).
Mapa del recorrido completo del clínico, derivado del código (`apps/web`) — no de la memoria. Un clínico, una sesión autenticada, **dos carriles de producto** desde la consola. **El carril B es el MVP**: la cuña del registro calificado por completitud, donde se construye el foso (*moat*). La extracción y el envío al registro son las dos bisagras que dependen de una bandera de configuración.
## El recorrido
[Sección titulada «El recorrido»](#el-recorrido)
```
flowchart TD
Landing["Landing · poktacare.com
capta al clínico"]:::xrepo --> SignIn
subgraph AUTH["Entrar"]
direction TB
SignIn["/iniciar-sesion
login Privy"]:::ship --> Me{"GET /api/users/me"}
Me -->|"403 access_required"| Req["/solicitar-acceso
none · pending · approved · rejected"]:::ship
Me -->|"sin onboarding"| Perfil["/consola/perfil
perfil + cédula"]:::ship
Me -->|"ok"| Home
Req -->|"aprobado · Entrar"| Home
Perfil --> Home
end
Home["/consola · ConsoleHome
inicio RheumAI"]:::ship
Home --> NC
Home --> Inv
subgraph LANEA["Carril A · Consulta → Plan de paciente"]
direction TB
NC["/consola/nueva-consulta"]:::ship --> AE["/consola/borradores/:id
Editor de aprobación"]:::ship
AE -->|"Compartir · magic link"| PP["/p/plan
plan del paciente (token, sin login)"]:::ship
end
subgraph LANEB["Carril B · Cuña BiobadamexAI — el MVP"]
direction TB
Up["/consola/biobadamex/subir
subir .docx/.doc/.pdf/.txt · ≤25"]:::ship
Wa["WhatsApp file-drop
consent → auth → acuse → ingesta"]:::gate
Up --> Inv
Wa --> Inv
Inv["/consola/biobadamex
inventario · filtros de estado + canal"]:::ship
Inv -->|"Extraer"| Ex{"extraer + calificar
QUEUE_ENABLED"}:::gate
Ex --> Rev["/consola/biobadamex/:id
Compuerta de revisión · checklist por variante"]:::ship
Rev --> Gr["/…/calificacion
la calificación — dato del foso"]:::ship
Rev -->|"Aprobar · dueño + checks"| Sub{"Enviar
solo dueño · SUBMIT_ENABLED · Charlson"}:::gate
Sub -->|"pg-boss → Tenki"| Reg[("Registro BIOBADAMEX")]:::ship
end
subgraph OUT["Un expediente · tres salidas"]
direction LR
O1["Completar
el expediente enviado"]:::ship
O2["Preguntar · Segunda opinión
rheumai.xyz — NO cableado aquí"]:::xrepo
O3["Plan de paciente
/p/plan · 'pregunta a RheumAI' pendiente"]:::stub
end
Reg --> O1
AE --> O3
Home -.->|"halo de profundidad"| O2
classDef ship fill:#e4eef8,stroke:#2b6cb0,stroke-width:1px,color:#1b3a57;
classDef gate fill:#f6ecd6,stroke:#9a6a10,stroke-width:1px,color:#5c3f08;
classDef stub fill:#e9ebef,stroke:#6b7280,stroke-width:1px,color:#3b3f48;
classDef xrepo fill:#ece6fb,stroke:#6d45d9,stroke-width:1px,color:#3d2680;
```
**Leyenda** (colores del diagrama): **azul** = enviado y accesible · **ámbar** = tras bandera / condicional · **gris** = pendiente / sin construir · **violeta** = app aparte (rheumai.xyz).
## Cada pantalla, cada punto de decisión
[Sección titulada «Cada pantalla, cada punto de decisión»](#cada-pantalla-cada-punto-de-decisión)
Los “puntos de decisión” son las bifurcaciones, compuertas y estados condicionales dentro de cada pantalla — donde el flujo se ramifica y donde tomamos decisiones de producto.
| Pantalla | Propósito | Puntos de decisión y compuertas | Estado |
| ----------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------- |
| `/iniciar-sesion` — Ingresar | Login alojado por Privy. | Al autenticar, retorna a `?next=` (p. ej. un enlace de revisión abierto desde WhatsApp aterriza ahí). | Enviado |
| `/solicitar-acceso` — Solicitar acceso | Compuerta *fail-closed*: un médico autenticado pero no aprobado pide acceso. | 4 estados de `useAccessStatus`: **none** → formulario (nombre + cédula + especialidad + consentimiento) · **pending** · **approved** → “Entrar” · **rejected**. Se renderiza en el shell lateral con logout. El correo se sella del lado servidor. | Enviado |
| `/consola/perfil` — Onboarding | Perfil de primera vez (nombre, especialidad, cédula). | Compuerta en `Shell`: \`canUseApp = onboarded | |
| `/consola` — Inicio | Inicio RheumAI — “Nueva consulta” + consultas recientes. | Estados carga / vacío / lista; filas con `draftId` enlazan al Editor de aprobación. Es la entrada del **Carril A** — el producto de consulta, distinto de la cuña. | Enviado |
| `/consola/biobadamex` — Inventario | Una lista para toda la cuña; el estado es un filtro, no una pantalla. | Filtro de estado (revisar → sin extraer → en proceso → aprobadas → enviadas → rechazadas → todas) con default inteligente. **Filtro de canal** ortogonal (whatsapp/consola/masiva) sólo aparece con >1 canal. “Extraer” en lote sólo en el filtro “sin extraer”. | Enviado |
| `/consola/biobadamex/subir` — Subir | Ingesta desde consola. | Dos modos: **pegar** (extrae ya) vs **archivo** (queda en inventario, se extrae después). Acepta `.docx .doc .pdf .txt`, ≤ 25 archivos. La extracción depende de `QUEUE_ENABLED`. | Enviado |
| WhatsApp file-drop — *(servidor, sin pantalla)* | El carril de ingesta nativo del chat. | Whitelist o *pass-through* a triage. consentimiento (ACEPTO/BAJA) → clasifica documento por extensión → **auth clínica fail-closed** → acuse exactly-once → ingesta (núcleo compartido) → hitos sin PHI. El coach es otra bandera (`COACH_ENABLED`, off por defecto). | Tras bandera |
| `/consola/biobadamex/:id` — Compuerta de revisión | Revisión de completitud + aprobación no evitable. No edita el expediente en silencio — sólo aprobar/rechazar/corregir. | **Checklist por variante**: AR → NAD/NAT/DAS-28/VSG/PCR; vasculitis → BVAS; LES → SLEDAI (no obligatorio, pendiente de dueño clínico); ES → Valentini/EUSTAR; EspA/Sjögren → sin instrumento forzado. **Compuerta de aprobación** = diagnóstico resuelto Y cada fila confirmada Y episodios de biológico confirmados. Corregir reinicia todas las confirmaciones. **Aquí no hay “Segunda opinión”.** | Enviado |
| `…/:id → Enviar` — Envío al registro | Empuja el expediente aprobado a BIOBADAMEX (pg-boss → Tenki). | Sólo en la pantalla terminal (aprobada). **Sólo dueño** (`draft.medicId === me.id`; los admin ven texto, sin botón) · compuerta dura `SUBMIT_ENABLED=1` · comorbilidades Charlson no documentadas exigen un checkbox de “confirmar ausencia”. | Tras bandera |
| `/…/calificacion` — La calificación | Primera impresión del médico — y la fuente del foso. El veredicto se muestra antes de cualquier pregunta. | Un *prompt* por campo faltante; cada uno resuelve a **acted** (→ compuerta de revisión) o **dismissed** (con motivo). Alimenta el registro de calidad de interrupción. | Enviado |
| `/consola/borradores/:id` — Editor de aprobación (Carril A) | La pantalla original de aprobación de consulta — diagnóstico + decisiones de tratamiento. | No evitable: todas las filas confirmadas antes de aprobar; códigos de baja confianza marcados. Ya aprobada, “Compartir” genera un magic link del paciente. | Enviado |
| `/p/plan` — Plan del paciente | Vista de sólo lectura por token (sin login) — el pago del Carril A. | Implementada completa (tratamiento vs lenguaje simple, procedencia). Un stub dentro: “Pregúntale a RheumAI · próximamente”. | Enviado (+ stub) |
## Dónde muerden las compuertas de lanzamiento
[Sección titulada «Dónde muerden las compuertas de lanzamiento»](#dónde-muerden-las-compuertas-de-lanzamiento)
Del registro de compuertas (`workstreams/biobadamexai/LAUNCH-GATES.md`), mapeadas sobre el flujo. **BLOCKER** = antes del lanzamiento público · **CDSS** = antes de anunciar “Preguntar” como apoyo a decisiones citado · **BUILD** = una superficie prometida aún sin cablear.
| Compuerta | Severidad | Dónde muerde |
| ------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------ | ---------------------------------------------------------------------------- |
| **G8 · G9** — orígenes Privy + smoke test de auth | Blocker | Paso de ingreso (login no funciona fuera de dominio hasta permitir orígenes) |
| **G4** — `GET /api/registry/queue` sin proteger | Blocker | rheumai.xyz (app aparte) |
| **G5 · G5c** — ingesta WhatsApp + coach | Blocker (ya **en producción**; la fila del registro está desactualizada) | Carril WhatsApp file-drop |
| **G5b** — bypass de auto-respuesta del triage de paciente | Blocker\* (aislado del piloto por la whitelist) | Ruta *pass-through* |
| **G1 · G2 · G3** — barandal recomienda-no-diagnostica, estadísticas sin verificar, unificar citas | CDSS | Salida “Preguntar” (se puede enviar acotado sin esto) |
| **G6** — cableado de Planes de paciente | Build | Salida Planes (el Carril A ya genera /p/plan; el chat in-plan es el stub) |
| **G10** — chequeo de deriva del diccionario de campos | Build | Compuerta de revisión |
| **G7** — bloque “Preguntar” en la landing (diseño) | Build | Landing |
## Decisiones por fijar
[Sección titulada «Decisiones por fijar»](#decisiones-por-fijar)
Lo que el mapa deja abierto para el flujo del MVP — las decisiones a tomar mientras lo recorremos.
1. **¿“Preguntar / Segunda opinión” está en el MVP, y desde dónde?** No está cableado en la consola — la compuerta de revisión no tiene UI de segunda opinión, y `consultRheumaSupport` es código muerto. La capacidad vive sólo en la app aparte rheumai.xyz. Decidir: enlazar hacia ella, cablearla, o sacarla del encuadre del MVP.
2. **¿Dos carriles o uno?** La consola lleva el producto original consulta→plan (Carril A) y la cuña BiobadamexAI (Carril B). El rebrand mantiene ambos como nav de nivel superior. Para el MVP, ¿el Carril A está en alcance, o el lanzamiento es estrictamente la cuña?
3. **¿Dónde aterriza *primero* el clínico — y sobre qué?** El inicio de consola es la superficie de consulta original, no la cuña. Para un clínico captado *para* BIOBADAMEX, ¿`/consola` debería abrir con la cuña (notas por revisar) en lugar de “Nueva consulta”?
4. **SLEDAI y los conjuntos obligatorios por variante.** Los instrumentos de LES se renderizan pero no son obligatorios (pendiente de dueño clínico); EspA/Sjögren no fuerzan instrumento. El conjunto de campos obligatorios por variante *es* la definición de completitud — necesita el visto bueno clínico antes de calificar notas reales.
5. **Envío al registro — ¿activo para el piloto?** El envío es sólo-dueño y con compuerta dura `SUBMIT_ENABLED`. Confirmar la postura de la bandera para el piloto de 20 médicos: ¿envían al registro en vivo, o paran en “aprobado” mientras se valida el ciclo?
# Repos y guías
> Directorio de enlaces — cada repo, su README/CONTRIBUTING, y dónde está desplegado.
## `pokta-grader-bioagent`
[Sección titulada «pokta-grader-bioagent»](#pokta-grader-bioagent)
El Grader — arm 3 del estudio, el producto en evolución.
* Repo: [github.com/poktalabs/pokta-grader-bioagent](https://github.com/poktalabs/pokta-grader-bioagent)
* Onboarding: [`CONTRIBUTING.md`](https://github.com/poktalabs/pokta-grader-bioagent/blob/main/CONTRIBUTING.md)
* Setup de entorno: [`dev/SETUP.md`](https://github.com/poktalabs/pokta-grader-bioagent/blob/main/dev/SETUP.md), [`dev/getting-started.md`](https://github.com/poktalabs/pokta-grader-bioagent/blob/main/dev/getting-started.md) (genéricos del framework BioAgents)
* Convenciones de ingeniería: [`AGENTS.md`](https://github.com/poktalabs/pokta-grader-bioagent/blob/main/AGENTS.md)
* Desplegado: Railway, `grader.poktacare.com`
* Runtime: Bun + Elysia
## `rheuma-ai-bioagent`
[Sección titulada «rheuma-ai-bioagent»](#rheuma-ai-bioagent)
RheumaAI — arm 2 del estudio, agente ya entrenado en reumatología. Origen del fork del Grader.
* Repo: [`poktalabs/rheuma-ai-bioagent`](https://github.com/poktalabs/rheuma-ai-bioagent) (privado)
* Onboarding: `README.md` (explica routes, tools, state model, LLM library)
* Desplegado: Railway, `rheumai.xyz`
* Runtime: Bun + Elysia — mismo stack que el Grader
## `pokta-care-monorepo`
[Sección titulada «pokta-care-monorepo»](#pokta-care-monorepo)
Orquestador del estudio (arm 1 en proceso, invoca los otros dos arms por HTTP, scoring, persistencia) y la app de producto real.
* Repo: [`poktalabs/pokta-care-monorepo`](https://github.com/poktalabs/pokta-care-monorepo)
* Scripts del benchmark: `apps/api/scripts/biobadamex-*.ts`
* Scoring: `apps/api/src/services/biobadamex-eval.ts`
* Invocación de arms: `apps/api/src/services/biobadamex-study-arms.ts` y `biobadamex-study-harness.ts`
* Desplegado: Vercel (`app.poktacare.com` / `api.poktacare.com`) + Railway
* Stack: TypeScript, Turbo, Drizzle, React
## Este sitio
[Sección titulada «Este sitio»](#este-sitio)
* Repo: [`poktalabs/pokta-care-docs`](https://github.com/poktalabs/pokta-care-docs)
* Stack: Astro + Starlight, desplegado en Cloudflare Pages bajo `docs.poktacare.com`
* Acceso: gateado — ver [Reglas](/guardrails/)
# Arquitectura de RheumAI
> Componentes, flujo de datos y puntos de integración del sistema actual.
## Mapa del sistema
[Sección titulada «Mapa del sistema»](#mapa-del-sistema)
```
flowchart TB
UI["Preact web app"] --> API["Elysia API"]
EXT["Cliente externo"] --> API
STUDY["BiobadamexAI study harness"] --> SR["Ruta stateless de estudio"]
API --> ACCESS["Auth · rate limits · x402 opcional"]
ACCESS --> SETUP["Conversación y estado"]
SETUP --> PLAN["PLANNING"]
PLAN --> TOOLS["Herramientas clínicas y fuentes"]
TOOLS --> RET["DAG / RAG / grafo"]
RET --> REPLY["REPLY o HYPOTHESIS"]
REPLY --> ORVS["ORVS + control PMID"]
ORVS --> ETH["Revisión ética"]
ETH --> UI
SR --> PLAN2["PLANNING con límite de herramientas"]
PLAN2 --> TOOLS2["Herramientas permitidas"]
TOOLS2 --> REPLY2["REPLY"]
REPLY2 --> STUDY
SETUP --> DB[("Supabase")]
TOOLS --> OBJ[("Storage S3 compatible")]
TOOLS --> LIT["PubMed · Semantic Scholar · OpenScholar · conocimiento local"]
```
## Componentes
[Sección titulada «Componentes»](#componentes)
### Cliente web
[Sección titulada «Cliente web»](#cliente-web)
La interfaz Preact gestiona sesiones, autenticación, carga de archivos, envío de mensajes, citas y estados de pago. El backend sirve el bundle y usa una ruta fallback para la aplicación de una sola página.
### API Elysia
[Sección titulada «API Elysia»](#api-elysia)
`src/index.ts` monta las rutas de autenticación, configuración, chat, investigación profunda, comunidad, búsqueda, criptografía y estudio. También sirve la aplicación web y expone endpoints de salud.
### Estado y persistencia
[Sección titulada «Estado y persistencia»](#estado-y-persistencia)
El chat normal crea usuarios, conversaciones, mensajes y estado en Supabase. Los archivos pueden ir a un proveedor S3 compatible y su texto analizado se agrega al estado del turno. Esto significa que `/api/chat` no es stateless.
La ruta de estudio usa UUID nulos y omite identificadores de mensaje y estado. Sus escrituras quedan desactivadas por diseño. Acepta texto desidentificado, puede fijar un proveedor/modelo y devuelve atribución, herramientas usadas, latencia y uso de tokens.
### Planificador y herramientas
[Sección titulada «Planificador y herramientas»](#planificador-y-herramientas)
El registro descubre carpetas bajo `src/tools/` y carga los `index.ts` habilitados. Variables `TOOL__ENABLED` pueden activar o desactivar una herramienta. El planificador decide qué fuentes consultar y cuál acción principal ejecutar.
Familias relevantes:
* Evidencia: PubMed, Semantic Scholar, OpenScholar, conocimiento local y grafo.
* Clínica: razonamiento clínico, interpretación de laboratorio, resumen para aseguradora y protocolos.
* Síntesis: `REPLY`, `HYPOTHESIS`, reflexión y revisión ética.
* Infraestructura: archivos, análisis de datos, almacenamiento y servicios x402.
Que una carpeta exista no garantiza que cargue. La carga es dinámica, registra errores y continúa con las herramientas disponibles.
### Recuperación y ranking
[Sección titulada «Recuperación y ranking»](#recuperación-y-ranking)
El flujo puede combinar recuperación por embeddings, grafo local y búsqueda semántica. Después, `DAG-RETRIEVAL` reordena documentos para acercar la evidencia a la consulta. Si no hay contexto útil, algunos modos regresan a una respuesta directa del modelo.
### Modelos
[Sección titulada «Modelos»](#modelos)
La biblioteca de LLM expone una interfaz común para varios proveedores. `resilientChatCompletion` implementa fallbacks. En el estudio, un contexto fijado puede obligar a que todos los saltos observables usen un mismo proveedor y modelo, con fallback deshabilitado.
## Dos flujos que no deben confundirse
[Sección titulada «Dos flujos que no deben confundirse»](#dos-flujos-que-no-deben-confundirse)
| Aspecto | Chat normal | Ruta de estudio |
| ----------------------------- | --------------------------------------------------- | ---------------------------------------------------- |
| Endpoint | `/api/chat` | `/v1/study/rheumaai-reply` |
| Persistencia | Sí | No, por diseño |
| Archivos | Sí | Solo texto de nota |
| Sesión | Conversación persistente | Una solicitud aislada |
| Auth | Sesión, cliente o x402 según configuración | Bearer dedicado; cerrado si falta la clave |
| Herramientas | Las elegidas por el planificador | Allowlist y techo de costo |
| ORVS posterior a la respuesta | Sí en el flujo normal; modos explícitos disponibles | No hay una llamada ORVS final explícita en esta ruta |
| Resultado | Respuesta conversacional | Respuesta, atribución y telemetría del turno |
## Dependencias externas
[Sección titulada «Dependencias externas»](#dependencias-externas)
* Supabase para conversaciones, mensajes, estados y búsquedas.
* Storage S3 compatible para archivos cuando está configurado.
* Proveedores LLM configurables.
* PubMed, Semantic Scholar, OpenScholar y corpus local para evidencia.
* Servicios médicos x402 opcionales.
* Railway y el dominio `rheumai.xyz` según la configuración del repositorio.
## Evidencia en código
[Sección titulada «Evidencia en código»](#evidencia-en-código)
* [Servidor y montaje de rutas](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/index.ts)
* [Ruta de chat](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/routes/chat.ts)
* [Orquestación de herramientas](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/services/chat/tools.ts)
* [Pipeline comparativo](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/services/chat/pipeline.ts)
* [Registro de herramientas](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/tools/index.ts)
* [Abstracción de LLM](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/llm/provider.ts)
* [Ruta de estudio](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/study/rheumaai-reply-route.ts)
# Evaluación de notas y completitud
> Qué puede hacer RheumAI hoy y cómo debe conectarse con el flujo controlado de BiobadamexAI.
> Esta página define el límite de RheumAI. La implementación actual de ingesta, completitud, revisión y envío vive en el [workflow de BiobadamexAI](/biobadamexai/).
## El problema
[Sección titulada «El problema»](#el-problema)
Una nota puede estar bien redactada y aun así omitir un dato necesario. También puede contener un dato correcto que el sistema no logra estructurar. Por eso hay que separar tres preguntas:
1. ¿Qué dice la nota?
2. ¿Qué dato exige el flujo clínico o el registro?
3. ¿Qué debe preguntar o corregir el médico antes de aprobar?
## Capacidad actual de RheumAI
[Sección titulada «Capacidad actual de RheumAI»](#capacidad-actual-de-rheumai)
RheumAI puede:
* recibir texto clínico y archivos en el chat;
* extraer texto de PDF, hojas de cálculo, CSV, JSON, Markdown y texto;
* usar visión para ciertos PDF e imágenes cuando está configurada;
* clasificar el tipo de consulta y detectar algunos términos de alerta;
* recuperar evidencia y producir una interpretación clínica;
* pedir datos faltantes mediante instrucciones de la persona y del prompt;
* revisar la calidad de su propia respuesta con ORVS;
* procesar una nota desidentificada por una ruta stateless para el estudio.
Estas capacidades ayudan a descubrir vacíos, pero hoy producen principalmente texto narrativo.
## Lo que no existe como contrato de RheumAI
[Sección titulada «Lo que no existe como contrato de RheumAI»](#lo-que-no-existe-como-contrato-de-rheumai)
El repositorio no expone todavía una respuesta estructurada con:
* cada campo requerido por BIOBADAMEX;
* valor extraído;
* fragmento de procedencia;
* estado `presente`, `ausente`, `desconocido` o `no_aplicable`;
* regla clínica que explica por qué el campo se exige;
* pregunta sugerida al médico;
* aprobación o rechazo del médico.
La ruta de estudio devuelve `content`, proveedor, modelo, herramientas, tokens, latencia y fallback. No devuelve una matriz de completitud de la nota. Tampoco aplica una comprobación determinista de desidentificación, límite específico para la longitud de la nota, ORVS final, revisión ética ni aprobación clínica. El llamador debe resolver esas obligaciones antes y después de la ruta.
## Diseño de integración recomendado
[Sección titulada «Diseño de integración recomendado»](#diseño-de-integración-recomendado)
El flujo institucional debe mantener responsabilidades separadas:
| Etapa | Responsable | Salida |
| --------------------- | ---------------------------------- | -------------------------------------------------------- |
| Ingesta | BiobadamexAI | Nota controlada y su identificador interno |
| Extracción | Servicio estructurado del registro | Campos candidatos con procedencia |
| Completitud | Reglas BIOBADAMEX por diagnóstico | Campos presentes, ausentes, desconocidos o no aplicables |
| Explicación clínica | RheumAI | Resumen de vacíos y preguntas útiles para el médico |
| Revisión de respuesta | ORVS y controles deterministas | Flags de exactitud, seguridad, citas y cobertura |
| Revisión humana | Médico autorizado | Corrección y aprobación explícita |
| Envío | Servicio controlado del registro | Solo datos aprobados, con lectura posterior |
## Regla central
[Sección titulada «Regla central»](#regla-central)
RheumAI debe **detectar y preguntar**, no inventar. Si la nota no contiene un dato, la integración debe conservar el estado ausente o desconocido según la regla clínica. Una inferencia del modelo nunca se convierte automáticamente en un valor del registro.
## Flujo para “completitud”
[Sección titulada «Flujo para “completitud”»](#flujo-para-completitud)
1. El extractor identifica valores candidatos y guarda el fragmento que los respalda.
2. El motor de reglas selecciona los campos aplicables al diagnóstico.
3. El sistema calcula completitud como presencia de campos aplicables. Exactitud y completitud se reportan por separado.
4. RheumAI transforma los vacíos en preguntas claras y priorizadas.
5. ORVS revisa la **respuesta de RheumAI**, no el expediente.
6. El médico confirma, corrige o marca que el dato no está disponible.
7. Solo la versión aprobada puede pasar al servicio de registro.
## Salvaguardas
[Sección titulada «Salvaguardas»](#salvaguardas)
* Usar notas sintéticas o desidentificadas fuera del entorno clínico autorizado.
* Recordar que el chat normal puede persistir mensajes y archivos. Los PDF largos que parecen artículos también pueden entrar al directorio local de conocimiento según las reglas actuales.
* No guardar notas reales en repositorios, documentación, logs ni fixtures.
* Mantener procedencia campo por campo.
* No convertir `0`, `false` o una fecha válida en “faltante” por una prueba booleana genérica.
* Mantener `ausente`, `desconocido` y `no aplicable` como estados distintos donde la regla clínica lo requiera.
* No permitir que ORVS o una respuesta narrativa salten la revisión del médico.
* Registrar qué modelo y versión de reglas produjo cada borrador.
## Criterios antes de conectar el flujo
[Sección titulada «Criterios antes de conectar el flujo»](#criterios-antes-de-conectar-el-flujo)
La integración no debe considerarse lista hasta que existan:
* contrato de entrada y salida versionado;
* reglas de completitud revisadas por diagnóstico;
* pruebas con notas sintéticas para campos presentes y faltantes;
* revisión clínica de las preguntas sugeridas;
* auditoría de que la ruta no escribe ni registra PHI de forma inesperada;
* control de autenticación servicio a servicio;
* fallback que preserve el borrador sin enviarlo si falla RheumAI.
## Evidencia en código
[Sección titulada «Evidencia en código»](#evidencia-en-código)
* [Ruta stateless de estudio](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/study/rheumaai-reply-route.ts)
* [Pruebas de la ruta de estudio](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/study/rheumaai-reply-route.test.ts)
* [Carga y análisis de archivos](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/tools/file-upload/index.ts)
* [Interpretación de laboratorio](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/tools/lab-interpretation/index.ts)
* [Razonamiento clínico](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/tools/clinical-reasoning/index.ts)
* [ORVS](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/tools/verification/index.ts)
# Integraciones y cambios seguros
> Contratos externos, configuración y pruebas mínimas antes de modificar RheumAI.
## Integraciones principales
[Sección titulada «Integraciones principales»](#integraciones-principales)
| Integración | Uso | Fallo esperado |
| ------------------------------ | ------------------------------------------------------- | ---------------------------------------------------------------------- |
| Supabase | Usuarios, conversaciones, mensajes, estados y búsquedas | El chat persistente puede fallar o degradarse |
| Storage S3 compatible | Archivos de conversación | El texto puede analizarse, pero la carga persistente puede omitirse |
| Proveedores LLM | Planificación, herramientas, respuesta y evaluación | `resilientChatCompletion` puede usar fallback fuera del estudio fijado |
| PubMed | Evidencia biomédica y PMID verificados | La respuesta debe declarar que no hubo PMID verificado |
| Semantic Scholar / OpenScholar | Literatura adicional | La herramienta registra el fallo y el pipeline puede continuar |
| Conocimiento local / grafo | Recuperación especializada | El pipeline puede usar otro método o una respuesta directa |
| BiobadamexAI study harness | Llama la ruta stateless con nota desidentificada | 401 sin bearer válido; 503 si la clave no está configurada |
| x402 | Pago opcional para API externa | 402 o 429 según el estado del pago y límites |
## Variables que definen comportamiento
[Sección titulada «Variables que definen comportamiento»](#variables-que-definen-comportamiento)
Esta lista nombra configuración, nunca valores:
* Servidor: `PORT`, `HOST`, `CORS_ORIGINS`.
* Estudio: `RHEUMAI_STUDY_API_KEY`, `CORS_STUDY_ORIGINS`, `RHEUMAI_STUDY_TOOL_ALLOWLIST`, `RHEUMAI_STUDY_MAX_PAID_CALLS`.
* Modelos: `REPLY_LLM_PROVIDER`, `REPLY_LLM_MODEL`, `HYP_LLM_PROVIDER`, `HYP_LLM_MODEL`, `STRUCTURED_LLM_MODEL`.
* ORVS: `ORVS_ENABLED`, `ORVS_VERSION`.
* Storage y base de datos: variables documentadas por los proveedores configurados; no copiarlas a la documentación.
* x402: `X402_ENABLED`, entorno, dirección de pago y credenciales del facilitador.
## Invariantes que no deben romperse
[Sección titulada «Invariantes que no deben romperse»](#invariantes-que-no-deben-romperse)
1. La ruta de estudio permanece stateless y cerrada si falta la clave.
2. El cuerpo de una nota no aparece en logs ni errores.
3. Un pin de estudio obliga a usar el proveedor/modelo fijado y desactiva fallback entre modelos.
4. Herramientas con costo quedan fuera del allowlist por defecto y el techo pagado predeterminado es cero.
5. Los PMID publicados deben provenir del conjunto recuperado en ese turno.
6. El chat clínico normal termina en `REPLY`, salvo flujos explícitos de protocolo o investigación.
7. Un fallo de RheumAI nunca autoriza el envío de datos al registro.
8. Solo un médico autoriza datos clínicos finales.
## Mapa de cambios y pruebas
[Sección titulada «Mapa de cambios y pruebas»](#mapa-de-cambios-y-pruebas)
| Si cambias | Revisa | Pruebas mínimas |
| -------------------------------- | ----------------------------------------------------------- | ------------------------------------------------------------------------------------------ |
| Montaje de rutas o CORS | `src/index.ts` y plugins Elysia | `bun run build`, smoke API sin mutaciones |
| Chat, orden de pasos o fallbacks | `src/routes/chat.ts`, `src/services/chat/*` | tests de herramientas y smoke con datos sintéticos |
| Ruta de estudio | `src/study/*` | `bun test src/study/rheumaai-reply-route.test.ts src/study/tool-cap.test.ts` |
| Proveedores/modelos | `src/llm/*` | `bun test src/llm/resilient.test.ts` y pruebas de pin |
| ORVS o citas | `src/tools/verification/*` y control PMID | tests unitarios, caso de PMID permitido y fabricado |
| Archivos o laboratorio | `src/tools/file-upload/*`, `src/tools/lab-interpretation/*` | `bun test src/tools/lab-interpretation/lab-interpretation.test.ts` con fixtures sintéticos |
| Herramientas dinámicas | `src/tools/index.ts` | arrancar el registro y revisar herramientas cargadas/fallidas |
| Interfaz | `client/src/*` | `bun run check`, `bun run build` |
## Verificación local recomendada
[Sección titulada «Verificación local recomendada»](#verificación-local-recomendada)
```bash
bun install --frozen-lockfile
bun run check
bun run build
bun test src/study/rheumaai-reply-route.test.ts \
src/study/tool-cap.test.ts \
src/tools/lab-interpretation/lab-interpretation.test.ts \
src/tools/knowledgeGraph/knowledgeGraph.test.ts
```
En el commit documentado, el conjunto enfocado anterior pasó **26 pruebas, 0 fallas** después de una instalación congelada. El arranque del registro también mostró warnings por herramientas sin credenciales o sin `index.ts`; las pruebas enfocadas no sustituyen una verificación completa del servidor.
Los smoke tests que llaman servicios externos requieren configuración válida. No deben ejecutarse contra producción ni con notas reales como parte de una modificación rutinaria.
## Estados del código
[Sección titulada «Estados del código»](#estados-del-código)
* **Activo:** importado por `src/index.ts`, la ruta de chat o una herramienta cargada y cubierto por pruebas observables.
* **Configurable:** activo solo cuando una variable o proveedor está presente.
* **Experimental:** modos comparativos, benchmarks y scripts no usados por defecto.
* **Placeholder o roto:** carpeta registrada sin `index.ts`, import que falla o integración que exige credenciales ausentes.
* **Legado:** conservado para compatibilidad, pero no es el camino preferido.
No documentes una carpeta como funcional solo porque existe. Confirma su importación, configuración y prueba.
La cobertura automática no incluye las rutas principales de chat, investigación profunda, autenticación, comunidad, Telegram o x402, ni tiene pruebas de cliente. Una modificación en esas áreas necesita smoke checks sintéticos además de los tests existentes.
## Evidencia en código
[Sección titulada «Evidencia en código»](#evidencia-en-código)
* [Rutas y CORS](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/index.ts)
* [Chat y controles PMID](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/routes/chat.ts)
* [Registro de herramientas](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/tools/index.ts)
* [Fallbacks de modelos](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/llm/resilient.ts)
* [Ruta de estudio](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/study/rheumaai-reply-route.ts)
* [Límite de herramientas del estudio](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/study/tool-cap.ts)
# ORVS en RheumAI
> Qué verifica ORVS, cómo se usa en el código actual y cuáles son sus límites.
ORVS significa **Optimistic Response Verification System**. El documento de referencia fue publicado como preprint por Erick Zamora el 28 de marzo de 2026 y tiene el DOI `10.55277/researchhub.qq2fdvw9`.\[1]
En RheumAI, ORVS es una revisión **posterior a la generación**. Evalúa la respuesta producida por el modelo. No valida directamente que la nota clínica de entrada tenga todos los campos que exige BIOBADAMEX.
## Preprint y código actual
[Sección titulada «Preprint y código actual»](#preprint-y-código-actual)
No son el mismo contrato. El preprint describe una rúbrica de cuatro dimensiones, hasta tres ciclos y escalamiento humano.\[1] El chat normal usa la versión heredada de seis dimensiones, intenta una regeneración y no implementa una cola de escalamiento humano ORVS. Los modos v2 son extensiones experimentales: Quick usa cuatro dimensiones y Deep usa ocho. Los pesos de Quick se parecen al preprint, pero el significado actual de `TMP` y `RSC` cambió a razonamiento temporal y completitud terapéutica.
## Flujo conceptual
[Sección titulada «Flujo conceptual»](#flujo-conceptual)
1. RheumAI genera una respuesta usando la consulta y la evidencia recuperada.
2. Un modelo evaluador puntúa varias dimensiones.
3. Si el resultado queda bajo el umbral, RheumAI agrega indicaciones de corrección y vuelve a generar o ampliar la respuesta.
4. Un control separado compara PMID citados contra los PMID recuperados en ese turno.
5. El sistema entrega la respuesta final o aplica un fallback si se agota el tiempo.
## Implementaciones presentes
[Sección titulada «Implementaciones presentes»](#implementaciones-presentes)
### Flujo normal de chat
[Sección titulada «Flujo normal de chat»](#flujo-normal-de-chat)
El chat normal llama a la verificación v1 para las consultas clínicas y para preguntas que no son saludos básicos. Evalúa seis dimensiones en escala 0–100:
* exactitud de citas;
* exactitud clínica;
* especificidad;
* alineación con la evidencia;
* completitud de la respuesta;
* ausencia de contradicciones.
El umbral actual es 70. Si no pasa, el chat intenta regenerar la respuesta usando los flags del evaluador. La respuesta regenerada no vuelve a pasar por ORVS en ese camino normal.
### Modos comparativos explícitos
[Sección titulada «Modos comparativos explícitos»](#modos-comparativos-explícitos)
Cuando el cliente envía `pipelineMode`, puede escoger `quick-orvs` o `full-orvs`:
| Modo | Dimensiones | Uso |
| ------------- | ---------------------------------------------------------------------------- | ------------------- |
| Quick ORVS v2 | exactitud clínica, seguridad, razonamiento temporal, completitud terapéutica | Revisión rápida |
| Deep ORVS v2 | las cuatro anteriores más completitud, citas, claridad y sesgo | Revisión más amplia |
La versión v2 calcula un compuesto 0–100 y puede regenerar o añadir una ampliación hasta dos veces.
## Controles de citas
[Sección titulada «Controles de citas»](#controles-de-citas)
Además del puntaje ORVS, el chat aplica una regla concreta para PMID:
* Si recuperó evidencia PubMed pero la respuesta no cita PMID, puede regenerar.
* Solo permite identificadores presentes en la evidencia recuperada en ese turno.
* Si persiste un PMID no permitido, lo reemplaza por `[PMID no verificado]`.
Este control es más determinista que pedir a un modelo que decida si una cita parece correcta.
## Límites importantes
[Sección titulada «Límites importantes»](#límites-importantes)
* ORVS evalúa una salida generada. No demuestra que una recomendación sea clínicamente correcta.
* El evaluador también es un modelo y puede equivocarse.
* El preprint describe una propuesta y experimentos retrospectivos con escenarios construidos.\[1] No demuestra resultados clínicos, aprobación regulatoria ni seguridad para uso autónomo.
* En el flujo actual, errores de parseo o API pueden producir un resultado marcado como `passed` con flags de error. Es un comportamiento fail-open, no una validación exitosa.
* `ORVS_ENABLED=false` desactiva la revisión.
* La ruta stateless del estudio no ejecuta una verificación ORVS final de forma explícita después de `REPLY`.
* La dimensión “completeness” de ORVS pregunta si la **respuesta cubre la consulta**. No es la completitud del expediente ni del registro.
## Relación con completitud de notas
[Sección titulada «Relación con completitud de notas»](#relación-con-completitud-de-notas)
Para BiobadamexAI se necesitan dos evaluaciones separadas:
1. **Completitud de la nota:** campos requeridos presentes, ausentes, desconocidos o no aplicables, según la enfermedad y el contexto del registro.
2. **Calidad de la respuesta:** exactitud, seguridad, evidencia, claridad y utilidad de lo que RheumAI comunica al médico.
ORVS puede cubrir la segunda. La primera necesita un esquema clínico explícito y procedencia campo por campo. Consulta [Evaluación de notas y completitud](/rheumai/clinical-notes/).
## Evidencia en código
[Sección titulada «Evidencia en código»](#evidencia-en-código)
* [Implementación ORVS](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/tools/verification/index.ts)
* [Prompts de evaluación](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/tools/verification/prompts.ts)
* [Uso en chat y control PMID](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/routes/chat.ts)
* [Modos comparativos](https://github.com/poktalabs/rheuma-ai-bioagent/blob/149508b3ab79da0b88c56deb63f6fb66e35d0e68/src/services/chat/pipeline.ts)
## Sources
[Sección titulada «Sources»](#sources)
\[1] — Optimistic Response Verification System > “Optimistic Response Verification System” > “March 28, 2026”
# Cómputo y transparencia de costos
> En qué corre el benchmark de evaluación de modelos, qué consume, y cómo controlamos el gasto — para partners de cómputo y patrocinadores.
Corremos un benchmark de LLMs de pesos abiertos para **extracción estructurada de datos clínicos** a partir de notas médicas de-identificadas — un estudio académico con un manuscrito en preparación. Esta página es un reporte transparente del **cómputo que consume el estudio**, los **factores de costo**, y los **controles** bajo los que corre. Está escrita para partners de cómputo y patrocinadores de investigación.
*Ninguna información que identifique a pacientes aparece en este benchmark ni en esta página; el corpus está de-identificado, y los detalles institucionales se omiten a propósito.*
## El benchmark, en un párrafo
[Sección titulada «El benchmark, en un párrafo»](#el-benchmark-en-un-párrafo)
Un **benchmark factorial** mide qué tan completa y correctamente distintos modelos extraen un conjunto fijo de campos clínicos estructurados de un corpus de **110 notas de-identificadas**, bajo **tres métodos de extracción** (“harnesses”). Corre enteramente sobre **modelos de pesos abiertos** servidos por [Nebius Token Factory](https://tokenfactory.nebius.com/).
| Harness | Descripción | Llamadas al modelo por nota |
| ---------- | --------------------------------------------------------------------------------------------------------------------- | --------------------------- |
| **plain** | Una sola llamada de extracción estructurada. | 1 |
| **agent** | Un agente clínico multi-paso (planeación → herramientas de retrieval → respuesta) cuya respuesta luego se estructura. | varias |
| **grader** | Un agente especializado de extracción/completeness. | varias |
**Diseño completo:** 4 modelos × 3 harnesses × 110 notas = **1,320 celdas de extracción**. Esto es investigación estándar de evaluación de modelos — sin tráfico de producción, sin generación masiva para redistribución. El resultado es un análisis comparativo de modelos abiertos para una tarea de NLP en salud, que la publicación resultante va a citar.
## Modelos bajo prueba
[Sección titulada «Modelos bajo prueba»](#modelos-bajo-prueba)
Cuatro modelos del catálogo de Nebius Token Factory:
| Modelo | Rol |
| ---------------------- | ---------------------------------- |
| `openai/gpt-oss-120b` | Línea base |
| `moonshotai/Kimi-K3` | Modelo de razonamiento bajo prueba |
| `zai-org/GLM-5.2` | Modelo de razonamiento bajo prueba |
| `MiniMaxAI/MiniMax-M3` | Modelo eficiente bajo prueba |
## Cómputo consumido a la fecha
[Sección titulada «Cómputo consumido a la fecha»](#cómputo-consumido-a-la-fecha)
Gasto medido en las corridas completas y parciales (cifras autoritativas del dashboard de uso del proveedor, 09 jul – 08 ago 2026):
| Modelo | Input (1M tokens) | Output (1M tokens) | Costo |
| ------------ | ----------------- | ------------------ | ---------- |
| gpt-oss-120b | 5.06 | 2.49 | $2.26 |
| Kimi-K3 | 1.63 | 1.34 | $25.06 |
| GLM-5.2 | 12.54 | 4.93 | $39.24 |
| MiniMax-M3 | 0.93 | 0.31 | $0.65 |
| **Total** | | | **$67.21** |
## Qué impulsa el costo
[Sección titulada «Qué impulsa el costo»](#qué-impulsa-el-costo)
1. **Los modelos de razonamiento pesado dominan.** GLM-5.2 ($39.24) y Kimi-K3 ($25.06) juntos son **\~95% del gasto**, casi enteramente en tokens de **output** (streams grandes de chain-of-thought / razonamiento por diseño). El modelo base más barato y el modelo eficiente fueron marginales en comparación. Este es en sí un hallazgo útil del benchmark: el volumen de output de un modelo de razonamiento, no el tamaño del input, es el costo real de una evaluación como esta.
2. **Los harnesses multi-paso hacen varias llamadas por nota.** Los métodos agent y grader corren una cadena de planeación → herramientas → respuesta → estructuración, así que una sola “celda” son varias llamadas facturables.
3. **Puesta en marcha de infraestructura.** El tuning inicial de concurrencia contra los rate limits produjo intentos reintentados antes de que estableciéramos el techo seguro — ya corregido (ver controles abajo).
## Para completar el benchmark
[Sección titulada «Para completar el benchmark»](#para-completar-el-benchmark)
**Pendiente:** una corrida limpia de los tres modelos no-base a través de los tres harnesses (\~990 celdas). Como el costo medido por celda de los modelos de razonamiento es alto, completar esto se estima en **\~$150–250**, dominado por GLM-5.2 y Kimi-K3. **Medimos el costo real por celda con un lote pequeño primero** y corremos el resto bajo presupuestos estrictos por oleada.
## Cómo controlamos el gasto
[Sección titulada «Cómo controlamos el gasto»](#cómo-controlamos-el-gasto)
* **Checkpoints humanos por oleada** — el benchmark corre un harness a la vez, y los resultados se revisan antes de gastar en la siguiente oleada. Nunca se gasta crédito sin un resultado que mostrar por él.
* **Pre-flight medido** — un lote pequeño establece el costo real por celda antes de cualquier corrida completa.
* **Concurrencia ajustada a los rate limits** — el techo de concurrencia seguro está medido, así que la API nunca se sobre-exige hasta reintentos.
## Equipo
[Sección titulada «Equipo»](#equipo)
* **Dr. Erick Zamora Tehozol** — Reumatología y validación clínica
* **Ing. Ángel Meléndez Córdoba** — Ingeniería
## Para patrocinadores y partners de cómputo
[Sección titulada «Para patrocinadores y partners de cómputo»](#para-patrocinadores-y-partners-de-cómputo)
Este benchmark es **uso académico y citable de un catálogo de modelos abiertos** para una publicación de NLP en salud. Si hospedas o financias inferencia de pesos abiertos y quieres que tus modelos estén representados en la comparación — o quieres patrocinar el cómputo que lo completa — el estudio es un uso transparente, bien instrumentado, y con destino de publicación de ese cómputo. Contáctanos a través de los medios en [poktacare.com](https://poktacare.com).
# Compute & Cost Transparency
> What the model-evaluation benchmark runs on, what it consumes, and how we control spend — for compute partners and sponsors.
We run an open-weight LLM benchmark for **structured clinical-data extraction** from de-identified physician notes — an academic study with a manuscript in preparation. This page is a transparent account of the **compute the study consumes**, the **cost drivers**, and the **controls** we run it under. It is written for compute partners and research sponsors.
*No patient-identifying information appears anywhere in this benchmark or on this page; the corpus is de-identified, and institutional details are omitted by design.*
## The benchmark, in one paragraph
[Section titled “The benchmark, in one paragraph”](#the-benchmark-in-one-paragraph)
A **factorial benchmark** measures how completely and accurately different models extract a fixed set of structured clinical fields from a corpus of **110 de-identified notes**, under **three extraction methods** (“harnesses”). It runs entirely on **open-weight models** served by [Nebius Token Factory](https://tokenfactory.nebius.com/).
| Harness | Description | Model calls per note |
| ---------- | ------------------------------------------------------------------------------------------------- | -------------------- |
| **plain** | A single structured-extraction call. | 1 |
| **agent** | A multi-step clinical agent (planning → retrieval tools → reply) whose answer is then structured. | several |
| **grader** | A specialized extraction/completeness agent. | several |
**Full design:** 4 models × 3 harnesses × 110 notes = **1,320 extraction cells**. This is standard model-evaluation research — no production traffic, no bulk generation for redistribution. The output is a comparative analysis of open models for a healthcare NLP task, which the resulting publication will cite.
## Models under test
[Section titled “Models under test”](#models-under-test)
Four models from the Nebius Token Factory catalog:
| Model | Role |
| ---------------------- | -------------------------- |
| `openai/gpt-oss-120b` | Baseline |
| `moonshotai/Kimi-K3` | Reasoning model under test |
| `zai-org/GLM-5.2` | Reasoning model under test |
| `MiniMaxAI/MiniMax-M3` | Efficient model under test |
## Compute footprint to date
[Section titled “Compute footprint to date”](#compute-footprint-to-date)
Measured spend on the completed and partial passes (authoritative figures from the provider’s usage dashboard, 09 Jul – 08 Aug 2026):
| Model | Input (1M tokens) | Output (1M tokens) | Cost |
| ------------ | ----------------- | ------------------ | ---------- |
| gpt-oss-120b | 5.06 | 2.49 | $2.26 |
| Kimi-K3 | 1.63 | 1.34 | $25.06 |
| GLM-5.2 | 12.54 | 4.93 | $39.24 |
| MiniMax-M3 | 0.93 | 0.31 | $0.65 |
| **Total** | | | **$67.21** |
## What drives the cost
[Section titled “What drives the cost”](#what-drives-the-cost)
1. **Reasoning-heavy models dominate.** GLM-5.2 ($39.24) and Kimi-K3 ($25.06) together are **\~95% of the spend**, almost entirely in **output** tokens (large chain-of-thought / reasoning streams by design). The cheaper baseline and the efficient model were negligible by comparison. This is itself a useful benchmark finding: reasoning-model output volume, not input size, is the real cost of an evaluation like this.
2. **Multi-step harnesses make several calls per note.** The agent and grader methods run a planning → tools → reply → structuring chain, so a single “cell” is multiple billable calls.
3. **Infrastructure bring-up.** Early concurrency tuning against rate limits produced retried attempts before we established the safe ceiling — now fixed (see controls below).
## To complete the benchmark
[Section titled “To complete the benchmark”](#to-complete-the-benchmark)
**Remaining:** a clean pass of the three non-baseline models across all three harnesses (\~990 cells). Because the reasoning models’ measured per-cell cost is high, completing this is estimated at **\~$150–250**, dominated by GLM-5.2 and Kimi-K3. We **meter the true per-cell cost on a small batch first** and run the remainder under strict per-wave budgets.
## How we control spend
[Section titled “How we control spend”](#how-we-control-spend)
* **Per-wave human checkpoints** — the benchmark runs one harness at a time, and results are reviewed before spending on the next wave. Credit is never spent without a result to show for it.
* **Metered pre-flight** — a small batch establishes the true per-cell cost before any full run.
* **Concurrency tuned to rate limits** — the safe concurrency ceiling is measured, so the API is never over-driven into retries.
## Team
[Section titled “Team”](#team)
* **Dr. Erick Zamora Tehozol** — Rheumatology & clinical validation
* **Ing. Ángel Meléndez Córdoba** — Engineering
## For sponsors and compute partners
[Section titled “For sponsors and compute partners”](#for-sponsors-and-compute-partners)
This benchmark is **citable academic use of an open-model catalog** for a healthcare-NLP publication. If you host or fund open-weight inference and want your models represented in the comparison — or want to sponsor the compute that completes it — the study is a transparent, well-instrumented, and publication-bound use of that compute. Reach out through the contacts on [poktacare.com](https://poktacare.com).