BioCreative

Data & API Inputs Map

The 28 data sources and agents feeding the Hub database (mjsgtszehjltxmbxtctz), grouped into 5 swimlanes: life science domain APIs, news and market signals, content and web-search learners, deep research plus the RAG/synthesis layer, and BioCreative's own operational signals (SEO/GEO, meetings, email, calendar). These are parallel, independent sources, not a sequential pipeline. Several feed directly into the Account & Contact Lifecycle diagram's Phase D signal loaders. Verified live against both VPS crontabs (bc-ops, bc-kb) and n8n on 2026-07-26 — not just skill-file claims.

Free (public API or internal SQL)
Paid (LLM classification, embeddings, or metered API)
Swimlane 1 — Life Science Domain APIs

Structured Public Data

7 systems, all free public APIs, forming the life-science backbone that also feeds account enrichment and the outreach signal loaders

Swimlane 2 — News & Market Intelligence

Live Signal Detection

Cross-feed correlation and news monitoring, synced to client spokes on a per-client cron schedule

Swimlane 3 — Content Learners & Web Search

Unstructured Knowledge Capture

Extract, classify, and catalog pattern applied to video, social, code, Reddit, RSS, and the open web

Swimlane 4 — Deep Research & Synthesis

Synthesis and Query Layer

Where everything above becomes queryable strategic intelligence, or gets synthesized into daily/weekly content proposals

Swimlane 5 — BioCreative's Own Operational Signals

Internal Performance & Communication Capture

Not life-science data — BC's own search performance, meetings, email, and calendar, all live-cron verified on bc-ops and bc-kb

Full Reference

What each source captures, what it costs, and where it lands

Swimlane 1 — Life Science Domain APIs
Swimlane 2 — News & Market Intelligence
Swimlane 3 — Content Learners & Web Search
Swimlane 4 — Deep Research & Synthesis
Swimlane 5 — BioCreative's Own Operational Signals