# NOSIBLE — Full Content Reference > NOSIBLE gives AI situational awareness through dated, enriched global events across politics, finance, and sport in near real time. --- --- title: "Know Everything, All The Time" description: "NOSIBLE gives AI situational awareness through dated, enriched global events across politics, finance, and sport in near real time." url: "https://nosible.com" --- §01 / 07 · NOSIBLE Live · Global Worldwide Web Surveillance # Know Everything, All The Time NOSIBLE gives AI complete situational awareness. Our 30‑year intelligence platform detects and enriches global events, from politics to finance to sports, in near‑real‑time. [START TRIAL](https://nosible.com/start-trial) [ENTER WORLD](https://nosible.world/world) [API REFERENCE](https://docs.nosible.com/) What we do ## Bottom Line Upfront? The NOSIBLE system 0 1 index 0 2 connect 0 3 intel index 031 01 ### We Index The Web We crawl the web without limits. We monitor every interest, in every geography and language. 02 ### We Connect Dots Our search engine connects similar documents through time creating a giant point-in-time network. 03 ### We Produce Intel AI discovers the events inside and files them into a deep ontology of genres, entities, and signals. Simulate our world ## Replay History NOSIBLE stores dated, ranked events so you can analyze any point in time. [Step into WORLD](https://nosible.world/world) Events 100M+ Sources 300K+ Languages 95 Countries 150+ Years Years/P.I.T 30 Events **100M+** Sources 300K+ Languages 95 Countries 150+ Years/P.I.T 30 Turn text into factors ## Quantify History NOSIBLE turns plain-language concepts into daily factors and stock betas. [Learn about semantic factors →](https://nosible.com/semantic-factors) 01 / Define ### Geopolitical risk War, terror, and conflict measured across dated world events. Anchors 24 Pairs 03 History 11Y 02 / Apply ### Company betas CVX +0.46% XOM +0.32% LMT +0.20% RTX +0.11% DAL +0.07% Daily semantic factor · 2015–2026 ![Daily geopolitical risk semantic factor from 2015 to 2026](https://nosible.com/data/macro-risk-indicators/plots/gpr_global-home-384.webp) 01 / Define + measure ### Geopolitical risk War, terror, and conflict measured across dated world events. Anchors 24 Polarity pairs 03 Daily history 11Y Daily semantic factor GPR_GLOBAL · 2015–2026 ![Daily geopolitical risk semantic factor from 2015 to 2026](https://nosible.com/data/macro-risk-indicators/plots/gpr_global.png) 02 / Apply ### Company betas Historical stock exposure after controlling for the market. Weekly semantic beta Market adjusted CVX Energy +0.46% XOM Energy +0.32% LMT Defense +0.20% RTX Defense +0.11% DAL Airline +0.07% Point-in-time verified Why we do it ## See Early Warnings NOSIBLE detects warning signs on the web before the related event occurs. [Read the ontology reference →](https://nosible.com/ontologies) Case study 01 Sugar Case studies 01 Sugar 02 Iran War 03 Theranos 04 Evergrande Timeline ### Sugar 2026 • 2027 Macroeconomic risk 1. 2026 · 02India's mills cut sugar estimates after satellite imagery shows yield damage. 2. 2026 · 04A below-normal monsoon is forecast, the first such April reading since 2015. 3. 2026 · 05India bans sugar exports with immediate effect. 4. 2027 · 05In six of seven comparable setups, sugar was materially higher a year later. The prediction **Sugar +86%** May 2027 Dithered sugarcane plants Built for scale ## Reliable Data APIs NOSIBLE APIs deliver web search and event data with predictable latency. [Read the API reference →](https://docs.nosible.com/) Discover the world ## Read Our Research Long-form on how we index, connect, and enrich the open web, plus the models behind NOSIBLE. [View all research](https://nosible.com/blog) [![High-resolution chart of selected 35B-A3B serving records across a 350-plus experiment campaign, from Qwen 3.5 FP8 on H200 at $0.218400 per million completion tokens in Experiment 006 to Qwen 3.6 source NVFP4 at $0.125959 in Experiment 332](https://nosible.com/images/2026/08/qwen-35b-a3b-serving-campaign-hero.svg) Lead research · Artificial Intelligence What 350+ Experiments Taught Us About Serving Qwen 3.6 35B-A3B We ran more than 350 35B-A3B serving experiments over five days across Qwen 3.5 and Qwen 3.6. The final Qwen 3.6 source-NVFP4 branch cut steady-state cost from $0.218400/M to $0.125959/M. Stuart Reid 2026-08-17 15 min read Read the research](https://nosible.com/blog/what-350-experiments-taught-us-about-serving-qwen-3-6) Latest notes 03 06 1. [![Thirty semantic factors rendered as daily time series on a dark grid, from geopolitical risk to firm-level political risk.](https://nosible.com/images/2026/08/semantic-factors-hero.png) Risk Indicator The Future That Could Have Been: Turning Web Text into Semantic Stock Betas 2026-08-05 · 11 min read](https://nosible.com/blog/the-future-that-could-have-been-turning-web-text-into-semantic-stock-betas) 2. [![NOSIBLE World knowledge graph showing entity connections over a decade](https://nosible.com/images/2026/07/kg-hero-decade.png) Web Search Point-in-Time Knowledge Graphs over Named Entities with NOSIBLE World 2026-07-16 · 8 min read](https://nosible.com/blog/point-in-time-knowledge-graphs-over-named-entities) 3. [![Signed contrast number line separating systemic and idiosyncratic risk](https://nosible.com/images/2026/06/walkthrough_step9_number_line.png) Artificial Intelligence Two Tricks for Turning Sentence Embeddings into Clean Features 2026-06-18 · 14 min read](https://nosible.com/blog/the-contrastive-geometry-of-risk) 4. [![Daily NOSIBLE Trade Policy Uncertainty index compared with the published index](https://nosible.com/images/2026/06/nosible-tpu-daily.png) Artificial Intelligence An Embedding-Based Approach to Trade and Economic Policy Uncertainty 2026-06-17 · 23 min read](https://nosible.com/blog/an-embedding-based-approach-to-trade-and-economic-policy-uncertainty) 5. [![S&P 500 news-stress overlay with drawdown and signal thresholds](https://nosible.com/images/2026/06/news-stress-overlay-hero.png) Trading Signals Turning News into a Risk-On/Risk-Off Equity Signal 2026-06-16 · 9 min read](https://nosible.com/blog/turning-news-into-a-risk-on-risk-off-equity-signal) 6. [![NOSIBLE geopolitical risk signal compared with published geopolitical risk indices](https://nosible.com/images/2026/06/nosible-gpr-vs-published.png) Risk Indicator We Rebuilt the Geopolitical Risk Index with Nosible World 2026-06-06 · 15 min read](https://nosible.com/blog/rebuilding-the-geopolitical-risk-index-from-nosible-world) S3 delivery, local API, or web API. ## Start your 90-day trial 90-day trial Review the diligence documents, schema, sample data, and delivery options. Then start your trial. [Start trial](https://nosible.com/start-trial) [Review docs](https://docs.nosible.com/) > NOSIBLE gives AI situational awareness through dated, enriched global events across politics, finance, and sport in near real time. **URL:** https://nosible.com --- --- title: "Request a free 90-day NOSIBLE World trial." description: "Request a free 90-day NOSIBLE World v1.2 trial and review the event schema, coverage analysis, samples, delivery options, DDQs, policies, and research." url: "https://nosible.com/start-trial" --- ### WORLD v1.2 trial evidence and coverage summary You can inspect the public April 2026 WORLD v1.2 sample before requesting a trial. It contains 900 ticker-linked events, selected as the 30 highest-coverage events for each day in the IPTC Economy, Business and Finance category. - 15,311,040 point-in-time corroborated events and 332,558,294 source records. - 38,097 tickers, 3,175,728 organizations, and 6,409,281 people. - Coverage views include geography, reporting languages, annual history, ticker coverage, named entities, seasonality, and ontology distributions. - WORLD v1.2 includes 15 ontology families with downloadable field guides and machine-readable releases. - 2025 averaged 1,439.9 ticker tags per day and 1,013.4 distinct tickers per day. The interactive event example is a complete Microsoft / Activision WORLD record from 18 January 2022 with 30 ranked provenance entries, 730 observed timestamps, source coverage across 2,392 domains, and a 3,072-dimension vector. Trial delivery options are S3 flat files, a local API, or the hosted Web API; due diligence, crawling, copyright, privacy, and security policies are available before a trial. STEP 01 OF 07 ## Trial request # Request a free 90-day NOSIBLE World trial. Tell us what your team wants to test. Full name Work email Organisation Organisation type Hedge fund Asset manager Proprietary trading firm Bank or broker-dealer Pension fund Sovereign wealth fund Family office Insurance company Data or technology company Corporate Government or public sector Academic or research institution Prefer not to disclose Functional role Portfolio management Quantitative research Fundamental research Data science or machine learning Data engineering or platform Data sourcing or procurement Legal or compliance Information security Product or strategy Executive management Prefer not to disclose Primary research domain Markets and corporate activity Macroeconomics and monetary policy Geopolitics and public policy Consumer and social trends Supply chains and trade Technology and cybersecurity Climate, energy, and sustainability Healthcare and life sciences Prefer not to disclose Product interests (optional) Worldwide coverage Ontology classification Point-in-time corroboration Cross-sectional coverage Named entity recognition Vectors and features Preferred trial delivery Not sure yet S3 Delivery Local API Web API Anything else we should know? (optional) Used only to arrange your trial. [Privacy policy](https://nosible.com/legal/privacy). Request Trial [NEXT Review the event and its fields](https://nosible.com/start-trial#data) STEP 02 OF 07 ## Event schema DATA DICTIONARY ### WORLD v1.2 fields Complete field names, types, definitions, and formats. ### Data Dictionary V1.2 Every WORLD v1.2 field, type, and definition. PDF / WORLD v1.2 / 2026-07-13 / 136.3 KiB [Download](https://nosible.com/trial/dictionaries/NOSIBLE%20World%20V1.2%20-%20Data%20Dictionary.pdf) [View the complete World data dictionary](https://nosible.com/data-dictionaries#world) API REFERENCE ### Search and World API docs Canonical endpoint schemas, authentication, parameters, responses, and integration examples. [Open API Reference ↗](https://docs.nosible.com/) ONTOLOGIES WORLD v1.2 includes 15 ontology families. The event below contains its assigned values and scores; the EDA Ontologies tab reports coverage and category distributions. [Explore all ontology field guides](https://nosible.com/ontologies) LOADING REAL WORLD EVENT Preparing the searchable event record. [NEXT Review coverage and EDA](https://nosible.com/start-trial#coverage) STEP 03 OF 07 ## Coverage and EDA WORLD V1.2 / COVERAGE REPORT [Download PDF](https://nosible.com/trial/coverage/nosible-world-v1.2-coverage-report.pdf) LOADING WORLD V1.2 EDA Preparing the coverage, panel, entity, and ontology views. [NEXT Review trial files and delivery](https://nosible.com/start-trial#files) STEP 04 OF 07 ## Files and delivery EVALUATION FILES ### Sample and join keys A month of ticker-linked events with the company and website reference files needed for joins. ### April 2026 Data Sample 900 ticker-linked events: 30 highest-coverage events daily throughout April 2026 in IPTC's Economy, Business and Finance category. The 6.0 MiB gzip expands to 46.9 MiB NDJSON; oai_vector is omitted. Gzipped NDJSON / WORLD v1.2 / April 2026 / 2026-07-12 / 6.0 MiB download / 46.9 MiB uncompressed [Download](https://nosible.com/trial/samples/nosible-world-v1.2-april-2026-ticker-linked-business-events.ndjson.gz) ### Ticker and company reference 38,097 WORLD v1.2 ticker mappings enriched with CompanyV4 reference metadata. XLSX / WORLD v1.2 / CompanyV4 / 2026-07-11 / 8.9 MB [Excel](https://nosible.com/trial/reference/tickers-companies-world-v1.2.xlsx) ### Website reference 113,787 WORLD v1.2 websites; 113,685 include WebsiteV10 metadata. XLSX / WORLD v1.2 / WebsiteV10 / 2026-07-11 / 6.2 MB [Excel](https://nosible.com/trial/reference/netlocs-websites-world-v1.2.xlsx) DELIVERY ### S3, Local API, or Web API **S3 Delivery** Flat files in S3. No executable software. Batch ingestion and full-history analysis. **Local API** Flat files with a Docker API in your environment. API access inside your own environment. [API Reference ↗](https://docs.nosible.com/) **Web API** NOSIBLE-hosted web API. No flat files. Hosted access without a local deployment. [API Reference ↗](https://docs.nosible.com/) IMPLEMENTATION ### Agent quickstarts Ready-to-use instructions for Claude, Codex, and Gemini. ### Claude Quickstart Bucket layout, reader contracts, validation, filtering, and semantic search for Claude. Markdown / WORLD v1.2 / 2026-07-12 / 18.9 KiB [Download](https://nosible.com/trial/guides/NOSIBLE%20World%20V1.2%20-%20Claude%20Guide.md) ### Codex Quickstart Agent instructions plus bucket layout, reader contracts, validation, filtering, and semantic search. Markdown / WORLD v1.2 / 2026-07-12 / 19.9 KiB [Download](https://nosible.com/trial/guides/NOSIBLE%20World%20V1.2%20-%20Codex%20Guide.md) ### Gemini Quickstart Agent instructions plus bucket layout, reader contracts, validation, filtering, and semantic search. Markdown / WORLD v1.2 / 2026-07-12 / 19.9 KiB [Download](https://nosible.com/trial/guides/NOSIBLE%20World%20V1.2%20-%20Gemini%20Guide.md) [NEXT Review DDQs and policies](https://nosible.com/start-trial#diligence) STEP 05 OF 07 ## DDQs and policies DDQS ### Due diligence questionnaires Completed responses for legal, compliance, procurement, and data-review teams. ### FISD DDQ NOSIBLE's completed responses to the 2024 FISD questionnaire. PDF / FISD 2024 / 2026-07-12 / 279.4 KiB [Download](https://nosible.com/trial/ddq/NOSIBLE%20World%20V1.2%20-%20FISD%20DDQ.pdf) POLICIES ### Crawling, copyright, privacy, and security [**Crawling** Collection from the public web and robots.txt. Read policy](https://nosible.com/legal/crawling) [**Copyright** Storage, display, attribution, and restrictions. Read policy](https://nosible.com/legal/copyright) [**Privacy** Personal data, retention, and removal. Read policy](https://nosible.com/legal/privacy) [**Security** Access controls and customer-data protection. Read policy](https://nosible.com/legal/security) [NEXT Review investment research](https://nosible.com/start-trial#research) STEP 06 OF 07 ## Investment Research RESEARCH & REPRODUCIBILITY ### Models and quantitative research Model weights, training data, methodology, benchmarks, and selected published results. SELECTED QUANTITATIVE RESEARCH [![NOSIBLE news stress signal and equity-market performance](https://nosible.com/images/2026/06/news-stress-overlay-hero.png) OUT-OF-SAMPLE STRATEGY **Turning News into a Risk-On/Risk-Off Equity Signal** Out-of-sample results from 2015-2026, plus Nasdaq and Russell 2000 tests. Read research](https://nosible.com/blog/turning-news-into-a-risk-on-risk-off-equity-signal) [![NOSIBLE daily Trade Policy Uncertainty index](https://nosible.com/images/2026/06/nosible-tpu-daily.png) PUBLISHED BENCHMARK **An Embedding-Based Approach to Trade and Economic Policy Uncertainty** TPU and EPU indices built from embeddings and compared with published benchmarks. Read research](https://nosible.com/blog/an-embedding-based-approach-to-trade-and-economic-policy-uncertainty) [![NOSIBLE and published Geopolitical Risk indices](https://nosible.com/images/2026/06/nosible-gpr-vs-published.png) PUBLISHED BENCHMARK **We Rebuilt the Geopolitical Risk Index with NOSIBLE World** A reconstructed GPR index with country, country-pair, and oil-risk breakdowns. Read research](https://nosible.com/blog/rebuilding-the-geopolitical-risk-index-from-nosible-world) [![Contrastive embedding score shown on a number line](https://nosible.com/images/2026/06/walkthrough_step9_number_line.png) FEATURE ENGINEERING **Two Tricks for Turning Sentence Embeddings into Clean Features** Two training-free embedding scores for relevance and systemic risk. Read research](https://nosible.com/blog/the-contrastive-geometry-of-risk) OPEN-SOURCE CLASSIFIERS FINANCIAL SENTIMENT #### Financial Sentiment v1.2 Base Labels financial impact as positive, neutral, or negative. [Open model](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) [Training dataset](https://huggingface.co/datasets/NOSIBLE/financial-sentiment) [Read methodology](https://nosible.com/blog/fast-enough-to-matter-productionizing-tiny-transformers-for-signal-extraction) FORWARD-LOOKING LANGUAGE #### Forward-Looking v1.2 Base Labels text as forward-looking or backward-looking. [Open model](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) [Training dataset](https://huggingface.co/datasets/NOSIBLE/forward-looking) [Read methodology](https://nosible.com/blog/fast-enough-to-matter-productionizing-tiny-transformers-for-signal-extraction) [NEXT Contact NOSIBLE](https://nosible.com/start-trial#contact) STEP 07 OF 07 ## Contact NOSIBLE [TECHNICAL CALENDAR **Technical Call** Compare S3 Delivery, Local API, and Web API. Open technical calendar](https://calendar.app.google/PZqw8fW9Wb8XqaFw5) [COMPLIANCE CALENDAR **Compliance Call** Ask about provenance, data rights, privacy, security, and controls. Open compliance calendar](https://calendar.app.google/PZqw8fW9Wb8XqaFw5) [COMMERCIAL CALENDAR **Commercial Call** Ask about trial terms, procurement, and production pricing. Open commercial calendar](https://calendar.app.google/PZqw8fW9Wb8XqaFw5) DIRECT EMAIL Stuart Reid · Chief Executive Officer [stuart@nosible.com](mailto:stuart@nosible.com?subject=NOSIBLE%20World%20trial) > Request a free 90-day NOSIBLE World v1.2 trial and review the event schema, coverage analysis, samples, delivery options, DDQs, policies, and research. **URL:** https://nosible.com/start-trial --- --- title: "Search the NOSIBLE website." description: "Search NOSIBLE products, research, comparisons, trial information and trust documentation." url: "https://nosible.com/search" --- NOSIBLE / Site search # Search the NOSIBLE website. Find product details, research, buyer comparisons, trial information and trust documentation. Search the NOSIBLE website Search ## Try a question How does NOSIBLE WORLD work? What is included in the free trial? How does NOSIBLE handle privacy and security? > Search NOSIBLE products, research, comparisons, trial information and trust documentation. **URL:** https://nosible.com/search --- --- title: "Media Kit" description: "Download official NOSIBLE logos, approved company copy, brand guidelines, illustrations and graphics, with media contact details for press and partners." url: "https://nosible.com/media-kit" --- Official media resources # Media Kit Copy, logos and visual assets for journalists, event teams and approved partners working on NOSIBLE coverage. [Download complete kit](https://nosible.com/media-kit/downloads/nosible-media-kit.zip) [Media contact](mailto:stuart@nosible.com?subject=NOSIBLE%20media%20request) Updated / 03 Aug 2026 01 / Copy ## Approved company copy Brand line Worldwide Web Surveillance Elevator pitch NOSIBLE gives AI complete situational awareness. Our 30-year intelligence platform detects and enriches global events, from politics to finance to sports, in near-real-time. [Download as text ↓](https://nosible.com/media-kit/approved-copy.txt) Click text to select · Ctrl/Cmd+C 02 / Logos ## NOSIBLE logos Use the primary logo unless the available space requires a more compact configuration. Complete packs include all 13 supplied files. ![Primary logo NOSIBLE logo](https://nosible.com/media-kit/library/logos/primary-logo-on-black.svg) Primary logo Default [Dark ↓](https://nosible.com/media-kit/library/logos/primary-logo-on-black.svg) [Transparent ↓](https://nosible.com/media-kit/library/logos/primary-logo-on-alpha.svg) [Light ↓](https://nosible.com/media-kit/library/logos/primary-logo-on-white.svg) ![Stacked logo NOSIBLE logo](https://nosible.com/media-kit/library/logos/stacked-logo-on-black.svg) Stacked logo Square or portrait [Dark ↓](https://nosible.com/media-kit/library/logos/stacked-logo-on-black.svg) [Transparent ↓](https://nosible.com/media-kit/library/logos/stacked-logo-on-alpha.svg) [Light ↓](https://nosible.com/media-kit/library/logos/stacked-logo-on-white.svg) ![Logomark NOSIBLE logo](https://nosible.com/media-kit/library/logos/logo-mark-green-on-black.svg) Logomark Avatar or icon [Green ↓](https://nosible.com/media-kit/library/logos/logo-mark-green-on-black.svg) [Dark ↓](https://nosible.com/media-kit/library/logos/logo-mark-on-black.svg) [Transparent ↓](https://nosible.com/media-kit/library/logos/logo-mark-on-alpha.svg) [Light ↓](https://nosible.com/media-kit/library/logos/logo-mark-on-white.svg) ![Logotype NOSIBLE logo](https://nosible.com/media-kit/library/logos/logo-type-on-black.svg) Logotype Type-led placement [Dark ↓](https://nosible.com/media-kit/library/logos/logo-type-on-black.svg) [Transparent ↓](https://nosible.com/media-kit/library/logos/logo-type-on-alpha.svg) [Light ↓](https://nosible.com/media-kit/library/logos/logo-type-on-white.svg) [SVG pack · 13 files ↓](https://nosible.com/media-kit/downloads/nosible-logo-pack-svg.zip) [PNG pack · 13 files ↓](https://nosible.com/media-kit/downloads/nosible-logo-pack-png.zip) [Usage guidelines ↓](https://nosible.com/media-kit/downloads/nosible-brand-bible.pdf) 03 / Visual system ## Type and colour A concise reference for the homepage palette and typographic hierarchy. Obsidian #0E0E10 Slate Core #1F1F23 Neo Green #21D081 ### Typographic hierarchy Use these four roles consistently to keep layouts clear and readable. Role Typeface Typical size Use Display Orbitron 72-112 px Page titles and major display copy. Heading Space Grotesk 36-52 px Section titles and short headlines. Body Inter 16-18 px Paragraphs, descriptions and captions. Label Space Mono 10-12 px Labels, filenames and technical details. 04 / Images ## Illustrations and graphics Download a complete image library or expand a collection to choose an individual file. ### Illustrations 21 high-resolution PNG files [Download all ↓](https://nosible.com/media-kit/downloads/nosible-illustrations-pack.zip) ### Graphic elements 20 numbered PNG files [Download all ↓](https://nosible.com/media-kit/downloads/nosible-graphics-pack.zip) Browse all 21 illustrations [![NOSIBLE illustration: Biker](https://nosible.com/media-kit/library/illustrations/biker.png) Biker PNG ↓](https://nosible.com/media-kit/library/illustrations/biker.png) [![NOSIBLE illustration: Biker 2](https://nosible.com/media-kit/library/illustrations/biker2.png) Biker 2 PNG ↓](https://nosible.com/media-kit/library/illustrations/biker2.png) [![NOSIBLE illustration: City](https://nosible.com/media-kit/library/illustrations/city.png) City PNG ↓](https://nosible.com/media-kit/library/illustrations/city.png) [![NOSIBLE illustration: Crow of People](https://nosible.com/media-kit/library/illustrations/crow-of-people.png) Crow of People PNG ↓](https://nosible.com/media-kit/library/illustrations/crow-of-people.png) [![NOSIBLE illustration: Cube](https://nosible.com/media-kit/library/illustrations/cube.png) Cube PNG ↓](https://nosible.com/media-kit/library/illustrations/cube.png) [![NOSIBLE illustration: Cyber](https://nosible.com/media-kit/library/illustrations/cyber.png) Cyber PNG ↓](https://nosible.com/media-kit/library/illustrations/cyber.png) [![NOSIBLE illustration: Determination](https://nosible.com/media-kit/library/illustrations/determination.png) Determination PNG ↓](https://nosible.com/media-kit/library/illustrations/determination.png) [![NOSIBLE illustration: Eye](https://nosible.com/media-kit/library/illustrations/eye.png) Eye PNG ↓](https://nosible.com/media-kit/library/illustrations/eye.png) [![NOSIBLE illustration: Future](https://nosible.com/media-kit/library/illustrations/future.png) Future PNG ↓](https://nosible.com/media-kit/library/illustrations/future.png) [![NOSIBLE illustration: Inspect](https://nosible.com/media-kit/library/illustrations/inspect.png) Inspect PNG ↓](https://nosible.com/media-kit/library/illustrations/inspect.png) [![NOSIBLE illustration: Late Nights](https://nosible.com/media-kit/library/illustrations/late-nights.png) Late Nights PNG ↓](https://nosible.com/media-kit/library/illustrations/late-nights.png) [![NOSIBLE illustration: Man and the City](https://nosible.com/media-kit/library/illustrations/man-and-the-city.png) Man and the City PNG ↓](https://nosible.com/media-kit/library/illustrations/man-and-the-city.png) [![NOSIBLE illustration: Pensive](https://nosible.com/media-kit/library/illustrations/pensive.png) Pensive PNG ↓](https://nosible.com/media-kit/library/illustrations/pensive.png) [![NOSIBLE illustration: Rage](https://nosible.com/media-kit/library/illustrations/rage.png) Rage PNG ↓](https://nosible.com/media-kit/library/illustrations/rage.png) [![NOSIBLE illustration: Spaceman](https://nosible.com/media-kit/library/illustrations/spaceman.png) Spaceman PNG ↓](https://nosible.com/media-kit/library/illustrations/spaceman.png) [![NOSIBLE illustration: Spacerace](https://nosible.com/media-kit/library/illustrations/spacerace.png) Spacerace PNG ↓](https://nosible.com/media-kit/library/illustrations/spacerace.png) [![NOSIBLE illustration: Sunrise](https://nosible.com/media-kit/library/illustrations/sunrise.png) Sunrise PNG ↓](https://nosible.com/media-kit/library/illustrations/sunrise.png) [![NOSIBLE illustration: The Sprinter](https://nosible.com/media-kit/library/illustrations/the-sprinter.png) The Sprinter PNG ↓](https://nosible.com/media-kit/library/illustrations/the-sprinter.png) [![NOSIBLE illustration: Track](https://nosible.com/media-kit/library/illustrations/track.png) Track PNG ↓](https://nosible.com/media-kit/library/illustrations/track.png) [![NOSIBLE illustration: Vortex](https://nosible.com/media-kit/library/illustrations/vortex.png) Vortex PNG ↓](https://nosible.com/media-kit/library/illustrations/vortex.png) [![NOSIBLE illustration: Wanderlust](https://nosible.com/media-kit/library/illustrations/wanderlust.png) Wanderlust PNG ↓](https://nosible.com/media-kit/library/illustrations/wanderlust.png) Browse all 20 graphic elements [![NOSIBLE graphic element: Graphic 01](https://nosible.com/media-kit/library/graphics/graphics-01.png) Graphic 01 PNG ↓](https://nosible.com/media-kit/library/graphics/graphics-01.png) [![NOSIBLE graphic element: Graphic 02](https://nosible.com/media-kit/library/graphics/graphics-02.png) Graphic 02 PNG ↓](https://nosible.com/media-kit/library/graphics/graphics-02.png) [![NOSIBLE graphic element: Graphic 03](https://nosible.com/media-kit/library/graphics/graphics-03.png) Graphic 03 PNG ↓](https://nosible.com/media-kit/library/graphics/graphics-03.png) [![NOSIBLE graphic element: Graphic 04](https://nosible.com/media-kit/library/graphics/graphics-04.png) Graphic 04 PNG ↓](https://nosible.com/media-kit/library/graphics/graphics-04.png) [![NOSIBLE graphic element: Graphic 05](https://nosible.com/media-kit/library/graphics/graphics-05.png) Graphic 05 PNG ↓](https://nosible.com/media-kit/library/graphics/graphics-05.png) [![NOSIBLE graphic element: Graphic 06](https://nosible.com/media-kit/library/graphics/graphics-06.png) Graphic 06 PNG ↓](https://nosible.com/media-kit/library/graphics/graphics-06.png) [![NOSIBLE graphic element: Graphic 07](https://nosible.com/media-kit/library/graphics/graphics-07.png) Graphic 07 PNG ↓](https://nosible.com/media-kit/library/graphics/graphics-07.png) [![NOSIBLE graphic element: Graphic 08](https://nosible.com/media-kit/library/graphics/graphics-08.png) Graphic 08 PNG ↓](https://nosible.com/media-kit/library/graphics/graphics-08.png) [![NOSIBLE graphic element: Graphic 09](https://nosible.com/media-kit/library/graphics/graphics-09.png) Graphic 09 PNG ↓](https://nosible.com/media-kit/library/graphics/graphics-09.png) [![NOSIBLE graphic element: Graphic 10](https://nosible.com/media-kit/library/graphics/graphics-10.png) Graphic 10 PNG ↓](https://nosible.com/media-kit/library/graphics/graphics-10.png) [![NOSIBLE graphic element: Graphic 11](https://nosible.com/media-kit/library/graphics/graphics-11.png) Graphic 11 PNG ↓](https://nosible.com/media-kit/library/graphics/graphics-11.png) [![NOSIBLE graphic element: Graphic 12](https://nosible.com/media-kit/library/graphics/graphics-12.png) Graphic 12 PNG ↓](https://nosible.com/media-kit/library/graphics/graphics-12.png) [![NOSIBLE graphic element: Graphic 13](https://nosible.com/media-kit/library/graphics/graphics-13.png) Graphic 13 PNG ↓](https://nosible.com/media-kit/library/graphics/graphics-13.png) [![NOSIBLE graphic element: Graphic 14](https://nosible.com/media-kit/library/graphics/graphics-14.png) Graphic 14 PNG ↓](https://nosible.com/media-kit/library/graphics/graphics-14.png) [![NOSIBLE graphic element: Graphic 15](https://nosible.com/media-kit/library/graphics/graphics-15.png) Graphic 15 PNG ↓](https://nosible.com/media-kit/library/graphics/graphics-15.png) [![NOSIBLE graphic element: Graphic 16](https://nosible.com/media-kit/library/graphics/graphics-16.png) Graphic 16 PNG ↓](https://nosible.com/media-kit/library/graphics/graphics-16.png) [![NOSIBLE graphic element: Graphic 17](https://nosible.com/media-kit/library/graphics/graphics-17.png) Graphic 17 PNG ↓](https://nosible.com/media-kit/library/graphics/graphics-17.png) [![NOSIBLE graphic element: Graphic 18](https://nosible.com/media-kit/library/graphics/graphics-18.png) Graphic 18 PNG ↓](https://nosible.com/media-kit/library/graphics/graphics-18.png) [![NOSIBLE graphic element: Graphic 19](https://nosible.com/media-kit/library/graphics/graphics-19.png) Graphic 19 PNG ↓](https://nosible.com/media-kit/library/graphics/graphics-19.png) [![NOSIBLE graphic element: Graphic 20](https://nosible.com/media-kit/library/graphics/graphics-20.png) Graphic 20 PNG ↓](https://nosible.com/media-kit/library/graphics/graphics-20.png) 05 / Downloads ## Downloads and contact Download the complete media kit, choose an individual pack, or contact NOSIBLE for help. ### Complete kit ZIP · 59.7 MB Copy, logos, guidelines, artwork and share cover [Download ↓](https://nosible.com/media-kit/downloads/nosible-media-kit.zip) ### Approved copy TXT Approved brand line and elevator pitch [Download ↓](https://nosible.com/media-kit/approved-copy.txt) ### Logo packs 2 ZIP files All 13 configurations in SVG and PNG [SVG pack ↓](https://nosible.com/media-kit/downloads/nosible-logo-pack-svg.zip) [PNG pack ↓](https://nosible.com/media-kit/downloads/nosible-logo-pack-png.zip) ### Illustrations ZIP · 21 files Complete high-resolution illustration folder [Download ↓](https://nosible.com/media-kit/downloads/nosible-illustrations-pack.zip) ### Graphic elements ZIP · 20 files Complete numbered graphics folder [Download ↓](https://nosible.com/media-kit/downloads/nosible-graphics-pack.zip) ### Visual guidelines PDF · 19.3 MB Supplied NOSIBLE brand guidelines [Download ↓](https://nosible.com/media-kit/downloads/nosible-brand-bible.pdf) [Media email stuart@nosible.com](mailto:stuart@nosible.com?subject=NOSIBLE%20media%20request) [Telephone +1 253 248 7193](tel:+12532487193) > Download official NOSIBLE logos, approved company copy, brand guidelines, illustrations and graphics, with media contact details for press and partners. **URL:** https://nosible.com/media-kit --- --- title: "API Reference" description: "Reference documentation for the NOSIBLE Search API and World API." url: "https://docs.nosible.com" last-modified: "2026-07-23" --- # API Reference The canonical API documentation for NOSIBLE Search API and World API. Use the API Reference for the current endpoint schemas, authentication, request parameters, response fields, examples, and error behavior. [Get an API key](https://app.nosible.com) ## World API World API provides structured, point-in-time event access. The documentation groups endpoints into GET EVENTS, SEARCH EVENTS, GET DETAILS, and HELPERS. GET EVENTS includes By day, By entity, By ticker, and By ontology. The By entity route is `GET /api/entities/events` and retrieves cursor-paged events for a canonical entity across a date window. It accepts a Bearer API key in the `Authorization` header, entity `type` and `name`, optional `limit`, opaque `cursor`, inclusive `from` and `to` dates, `order`, `include`, `include_vector`, and `include_live` parameters. World API responses can include a stable `schema`, canonical `entity`, `total`, `count`, `events`, `next_cursor`, `date_window`, `live_events`, `hydration_misses`, `as_of`, and `took_ms`. Send `next_cursor` back unchanged for the next page; do not infer a cursor from offsets or dates. World API errors documented at the reference include invalid requests, access denied, missing resources, expired cursors, rate limiting, and backend configuration, error, or timeout responses. Respect `Retry-After` when it is supplied for a 429 response. ## Search API Search API exposes web retrieval and scraping endpoints with an `Api-Key` header. The documented families are Web Search, Web Agents, Web Scraper, Web Mentions, Web Alerts, and Helpers. The Search API page on nosible.com covers `POST /search/v2/fast-search`, `POST /search/v2/bulk-search`, `POST /search/v2/rich-search`, `POST /search/v2/scrape-url`, and `POST /search/v2/search`. The canonical API Reference should be used for the complete, current parameter and response contract. ## Documentation guidance Questions about endpoint paths, authentication, parameters, pagination, response shapes, errors, rate limits, SDK usage, or the difference between Search API and World API should be answered from this source when it is retrieved. The canonical documentation is the API Reference at [docs.nosible.com](https://docs.nosible.com). **URL:** https://docs.nosible.com --- --- title: "Turn web text into semantic stock betas" description: "An inspectable catalogue of 30 open-source semantic factors for quant portfolio managers: 24 relevance anchors, 3 polarity pairs, daily history, code, and CSV." url: "https://nosible.com/semantic-factors" --- semantic-factors / open source # Turn web text into semantic stock betas semantic-factors turns NOSIBLE World text into daily semantic factors. Define a concept in sentences, choose polarity, and export the factor for research. [Explore 30 semantic factors](https://nosible.com/semantic-factors#explorer) [View the Python package](https://github.com/NosibleAI/semantic-factors) Recompute access: NOSIBLE API key + OpenRouter API key What semantic-factors does? 01 Define Write relevance anchors and polarity pairs in plain text. 02 Score Use NOSIBLE World and OpenRouter to score the definitions. 03 Export Inspect the curves, code, and CSV in your own workflow. STEP 1: WEB TEXT TO RISK ## Turn text on the web into a signal Use sentences to turn web text into a daily semantic factor you can inspect, export, and rerun. Select a semantic factor 01 - Geopolitical risk 02 - Trade policy uncertainty 03 - Pandemic and health risk 04 - Sanctions and export controls 05 - Recession nowcast 06 - Oil-supply risk 07 - Supply-chain pressure 08 - Economic policy uncertainty 09 - Financial stress 10 - Energy security 11 - Inflation attention 12 - Climate-policy uncertainty 13 - Bank-regulation uncertainty 14 - Fiscal uncertainty 15 - Country geopolitical risk 16 - Bilateral tension 17 - Nuclear threat 18 - Food security 19 - Commodity supply risk 20 - Sovereign distress 21 - Risk-on / risk-off 22 - Migration policy 23 - Currency crisis 24 - Cyber risk 25 - Terrorism and unrest 26 - Partisan conflict 27 - Monetary policy uncertainty 28 - Equity volatility tracker 29 - Firm-level political risk 30 - Hawk-dove sentiment 01 Geopolitical risk 02 Trade policy uncertainty 03 Pandemic and health risk 04 Sanctions and export controls 05 Recession nowcast 06 Oil-supply risk 07 Supply-chain pressure 08 Economic policy uncertainty 09 Financial stress 10 Energy security 11 Inflation attention 12 Climate-policy uncertainty 13 Bank-regulation uncertainty 14 Fiscal uncertainty 15 Country geopolitical risk 16 Bilateral tension 17 Nuclear threat 18 Food security 19 Commodity supply risk 20 Sovereign distress 21 Risk-on / risk-off 22 Migration policy 23 Currency crisis 24 Cyber risk 25 Terrorism and unrest 26 Partisan conflict 27 Monetary policy uncertainty 28 Equity volatility tracker 29 Firm-level political risk 30 Hawk-dove sentiment ### Geopolitical risk War, terror, and conflict as a share of world coverage. Chart Anchors Polarity Code Research ![Geopolitical risk semantic factor plot with raw daily observations and a geometric smoothing curve](https://nosible.com/data/semantic-factors/plots/gpr_global.png?v=2026-08-05T13%3A21%3A16.5947698Z-2e838bb84574) Click to expand Public snapshot through 2026-06-26 4,106 observed days 2.1 % unavailable / not observed [Download chart](https://nosible.com/data/semantic-factors/plots/gpr_global.png?v=2026-08-05T13%3A21%3A16.5947698Z-2e838bb84574) [Download data CSV](https://nosible.com/data/semantic-factors/gpr_global.csv?v=2026-08-05T13%3A21%3A16.5947698Z-2e838bb84574) Current scope global Historical output nos_dt Denominator global Weighting breadth Direction non-negative breadth Filter none STEP 2: SIGNAL TO BETAS ## Turn the signal into stock betas Use aligned returns to turn that signal into stock betas you can compare, explain, and use directly. 01 / SERIES ### Start with F t Choose one semantic factor and keep its scope, threshold, and history fixed across one clearly defined measurement window. date F t 02 / ALIGN ### Match dates Put the factor and adjusted closing prices on the same weekly clock, then calculate matching weekly returns and factor changes. factor change stock return week 01 week 02 week 03 week 04 03 / REGRESS ### Estimate beta Regress the stock return on the market return and the change in the semantic factor over the same weekly window for each company. factor change 04 / READ ### Explain the number A beta is a conditional historical association, not a claim that the factor caused the return or that the relationship will persist. LMT.US beta **+0.20%** Estimated weekly response to a one-standard-deviation GPR rise after market control. The regression `r i,t = α i + β market,i r market,t + β semantic,i ΔF t + ε i,t` Use the factor change when the question is sensitivity to a new shock. Standardize that change so the semantic beta reads as the stock return response to one typical factor move. Add sector, country, or other controls when the research question requires them. What the sign means Positive The stock tended to rise when the factor rose, after the market control. Near zero The sample does not show a strong incremental relationship to this factor. Negative The stock tended to fall when the factor rose, after the market control. Ten-stock example ### The same stocks, different exposures Choose a semantic factor Geopolitical risk Country geopolitical risk Bilateral tension Oil-supply risk Sanctions and export controls Nuclear threat Cyber risk Terrorism and unrest Economic policy uncertainty Trade policy uncertainty Monetary policy uncertainty Fiscal uncertainty Bank-regulation uncertainty Climate-policy uncertainty Partisan conflict Migration policy Financial stress Risk-on / risk-off Recession nowcast Inflation attention Equity volatility tracker Hawk-dove sentiment Commodity supply risk Energy security Food security Supply-chain pressure Pandemic and health risk Sovereign distress Currency crisis Firm-level political risk Geopolitical risk 2015-03-30 to 2026-06-22 586 weekly observations RTX.US defense +0.11% LMT.US defense +0.20% XOM.US energy +0.32% CVX.US energy +0.46% CAT.US industrial +0.03% DAL.US airline +0.07% AAPL.US technology +0.04% NVDA.US semiconductors +0.03% JPM.US financials +0.04% WMT.US retail -0.03% Semantic beta is the stock return response to a one-standard-deviation weekly change in the 30-day log1p geometric mean of the 0.25-0.50 k=6 tanh-squashed daily semantic factor after controlling for the weekly SPY.US return. The chart adds the 7-day arithmetic average of the same 30-day geometric mean for presentation. The values are full-sample OLS estimates from EODHD adjusted_close adjusted closes, not a backtest or recommendation. Build your next semantic factor ## Build your own semantic factor today! Get World access to recompute the examples, download all 30 semantic factors, or read the implementation on GitHub. [Start a trial](https://nosible.com/start-trial) [Download all 30 geometric series CSV](https://nosible.com/data/semantic-factors/nosible-macro-risk-indicators-daily.csv?v=2026-08-05T13%3A21%3A16.5947698Z-2e838bb84574) [Visit GitHub](https://github.com/NosibleAI/semantic-factors) Reference library ## Inspect the classifications behind semantic factors [Asset Class Ontology Map text-derived shocks to equity, credit, rates, commodity, currency, and derivative exposures. Open reference →](https://nosible.com/ontologies/asset-classes) [GICS Industry Classification Connect company-level semantic betas to consistent sectors and industries. Open reference →](https://nosible.com/ontologies/gics) [PLOVER Political Events Separate geopolitical threats, sanctions, protests, cooperation, and material conflict. Open reference →](https://nosible.com/ontologies/plover) > An inspectable catalogue of 30 open-source semantic factors for quant portfolio managers: 24 relevance anchors, 3 polarity pairs, daily history, code, and CSV. **URL:** https://nosible.com/semantic-factors --- --- title: "Our / Research" description: "Original NOSIBLE research on point-in-time data, search, embeddings, market signals and AI-native media intelligence." url: "https://nosible.com/blog" --- /res Research # Our / Research [Ontology field guides Definitions, category evidence, example usages, and downloadable counts.](https://nosible.com/ontologies) [Data dictionaries Field-level references for World and every Search API endpoint.](https://nosible.com/data-dictionaries) ![High-resolution chart of selected 35B-A3B serving records across a 350-plus experiment campaign, from Qwen 3.5 FP8 on H200 at $0.218400 per million completion tokens in Experiment 006 to Qwen 3.6 source NVFP4 at $0.125959 in Experiment 332](https://nosible.com/images/2026/08/qwen-35b-a3b-serving-campaign-hero.svg) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) ## [What 350+ Experiments Taught Us About Serving Qwen 3.6 35B-A3B](https://nosible.com/blog/what-350-experiments-taught-us-about-serving-qwen-3-6) We ran more than 350 35B-A3B serving experiments over five days across Qwen 3.5 and Qwen 3.6. The final Qwen 3.6 source-NVFP4 branch cut steady-state cost from $0.218400/M to $0.125959/M. 2026-08-17 15 min read ![Thirty semantic factors rendered as daily time series on a dark grid, from geopolitical risk to firm-level political risk.](https://nosible.com/images/2026/08/semantic-factors-hero.png) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ## [The Future That Could Have Been: Turning Web Text into Semantic Stock Betas](https://nosible.com/blog/the-future-that-could-have-been-turning-web-text-into-semantic-stock-betas) Turn web text into transparent daily risk factors and stock-specific betas with an open-source, reproducible Python workflow. 2026-08-05 11 min read ![NOSIBLE World knowledge graph showing entity connections over a decade](https://nosible.com/images/2026/07/kg-hero-decade.png) [Web Search](https://nosible.com/blog/tag/web-search) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ## [Point-in-Time Knowledge Graphs over Named Entities with NOSIBLE World](https://nosible.com/blog/point-in-time-knowledge-graphs-over-named-entities) Companies change: people leave, products die, mergers happen. This post shows how to build point-in-time knowledge graphs over any company from news alone, using NOSIBLE World, which covers 38 thousand tickers, 3.2 million organizations and 6.4 million people. Named entity recognition supplies the nodes, lift-scored co-mention supplies the links, and the document date keeps every yearly view safe to backtest. 2026-07-16 8 min read ![Signed contrast number line separating systemic and idiosyncratic risk](https://nosible.com/images/2026/06/walkthrough_step9_number_line.png) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ## [Two Tricks for Turning Sentence Embeddings into Clean Features](https://nosible.com/blog/the-contrastive-geometry-of-risk) A ten-step, training-free walkthrough that turns a frozen OpenAI text embedding into clean classifications: a multiclass relevance score sorts events into local, national, and global buckets, and a contrastive binary score splits systemic from idiosyncratic risk. Verified on real warnings from NOSIBLE World, the geometry matches Google's gemini-2.5-flash while staying deterministic, auditable, and effectively free. 2026-06-18 14 min read ![Daily NOSIBLE Trade Policy Uncertainty index compared with the published index](https://nosible.com/images/2026/06/nosible-tpu-daily.png) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ## [An Embedding-Based Approach to Trade and Economic Policy Uncertainty](https://nosible.com/blog/an-embedding-based-approach-to-trade-and-economic-policy-uncertainty) The Fed's Trade Policy Uncertainty index counts keywords across seven newspapers. We rebuilt it from 14.9 million NOSIBLE World events using only embeddings and five sentences, no keywords. It matches the published benchmark at 0.87 on monthly levels and 0.82 on monthly changes, as closely as the two official versions match each other. The same method, extended to sixty sentences, rebuilds the broader Economic Policy Uncertainty index and its national-security and healthcare categories. 2026-06-17 23 min read ![S&P 500 news-stress overlay with drawdown and signal thresholds](https://nosible.com/images/2026/06/news-stress-overlay-hero.png) [Trading Signals](https://nosible.com/blog/tag/trading-signals) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ## [Turning News into a Risk-On/Risk-Off Equity Signal](https://nosible.com/blog/turning-news-into-a-risk-on-risk-off-equity-signal) We built a risk-on/risk-off trading signal from the NOSIBLE event database that measures how much of the global news flow is about market-stress themes, holding equities when that reading is low and moving to T-bills when it spikes. Selected on 2010 to 2013 and tested on an untouched 2015 to 2026 window, it held the S&P 500's buy-and-hold return (+254% versus +269%) while cutting the maximum drawdown from −34% to −18% and raising the Sharpe ratio from 0.64 to 0.89. The same rule transfers unchanged to the Nasdaq and the Russell 2000. 2026-06-16 9 min read ![NOSIBLE geopolitical risk signal compared with published geopolitical risk indices](https://nosible.com/images/2026/06/nosible-gpr-vs-published.png) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ## [We Rebuilt the Geopolitical Risk Index with Nosible World](https://nosible.com/blog/rebuilding-the-geopolitical-risk-index-from-nosible-world) Markets move on geopolitics, but risk models cannot read the news. We turned 13.2 million news events into a geopolitical risk signal, matched the Federal Reserve benchmark, and broke it down by country, by country pair, and into an oil supply-risk signal. 2026-06-06 15 min read ![Running sprinter illustration representing efficient financial-sentiment model training](https://nosible.com/blog/illustrations/the-sprinter.png) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [Financial Sentiment](https://nosible.com/blog/tag/financial-sentiment) ## [Matching GPT-5.1 at Financial Sentiment with Active Learning and Qwen3](https://nosible.com/blog/fast-enough-to-matter-productionizing-tiny-transformers-for-signal-extraction) Here's how we fine-tuned Qwen3 0.6B to beat FinBERT and match GPT-5.1 accuracy. Complete with open-source models, datasets, and training scripts. Spoiler alert: active learning is all you need. 2025-12-12 27 min read ![Abstract vortex illustration representing self-organizing web-scale search facets](https://nosible.com/blog/illustrations/vortex.png) [Web Search](https://nosible.com/blog/tag/web-search) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) ## [Can Faceted Search at Web-Scale Self Organize?](https://nosible.com/blog/can-faceted-search-at-web-scale-self-organize) Can Faceted Search at Web-Scale Self Organize? As it turns out, yes it can! In this post we outline our new and improved adaptive named entity tagging system! 2025-10-16 7 min read ![Cybernaut-1 illustration representing agentic search with Monte Carlo Tree Search](https://nosible.com/blog/illustrations/cyber.png) [Web Search](https://nosible.com/blog/tag/web-search) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) ## [Introducing Cybernaut-1: Agentic Search using MCTS](https://nosible.com/blog/introducing-cybernaut-1-agentic-search-with-mcts) Cybernaut-1 combines our powerful hybrid-3 search algorithm with LLM-guided Monte Carlo Tree Search to deliver world class search results on difficult queries. 2025-08-26 2 min read ![Railway-track illustration representing the road to Cybernaut-1](https://nosible.com/blog/illustrations/track.png) [Web Search](https://nosible.com/blog/tag/web-search) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) ## [The Road to Cybernaut-1: Rebuilding Search for AI](https://nosible.com/blog/the-road-to-cybernaut-1) AI needs its own search engine. This is how we’re rebuilding search for AI -- and the road to Cybernaut-1, the first high-trust agentic search engine. 2025-08-20 17 min read ![3D cube illustration representing LLM ensemble distillation](https://nosible.com/blog/illustrations/cube.png) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [Financial Sentiment](https://nosible.com/blog/tag/financial-sentiment) ## [A Pattern for Scaling the Value Proposition of LLMs: Ensemble and Distil 🚀](https://nosible.com/blog/ensemble-and-distil) We introduce the ensemble and distil data pattern and use it to fit an ordinary least squares linear regression that outperforms GPT-4 at financial news sentiment classification using sentence transformer embeddings as features. 2024-02-06 12 min read ![Abstract eye illustration representing financial-news sentiment analysis](https://nosible.com/blog/illustrations/eye.png) [Financial Sentiment](https://nosible.com/blog/tag/financial-sentiment) ## [News Sentiment Showdown: Who Checks Vibes Best?](https://nosible.com/blog/news-sentiment-showdown-who-checks-vibes-best) A comparison of sentiment classifications made by TextBlob, VADER, Flair, SigmaFSA, FinBERT, FinBERT-Tone, Text-Bison, Text-Unicorn, Gemini-Pro, GPT-3.5, GPT-4, and GPT-4-Turbo. We look at accuracy, time, and cost and include a dataset of 10,368 labelled news stories (with code) for our followers. 2024-01-28 12 min read ![Magnifying-glass illustration representing vector search across company news](https://nosible.com/blog/illustrations/inspect.png) [Web Search](https://nosible.com/blog/tag/web-search) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) [Trading Signals](https://nosible.com/blog/tag/trading-signals) ## [Using Vector Search to See Signals in Company News](https://nosible.com/blog/using-vector-search-to-see-signals-in-company-news) How we use vector search to extract investment signals from a multi-terabyte company news dataset that currently contains over 55 million embeddings, 150+ million sentences, 4+ billion words, and 5+ billion GPT tokens. 2024-01-21 21 min read > Original NOSIBLE research on point-in-time data, search, embeddings, market signals and AI-native media intelligence. **URL:** https://nosible.com/blog --- --- title: "Artificial Intelligence" description: "Research on AI systems, agents, and language-model applications for search and quantitative research." url: "https://nosible.com/blog/tag/artificial-intelligence" --- /res Research # Artificial Intelligence Research on AI systems, agents, and language-model applications for search and quantitative research. [All](https://nosible.com/blog) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Financial Sentiment](https://nosible.com/blog/tag/financial-sentiment) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [Trading Signals](https://nosible.com/blog/tag/trading-signals) [Web Search](https://nosible.com/blog/tag/web-search) 8 articles ![High-resolution chart of selected 35B-A3B serving records across a 350-plus experiment campaign, from Qwen 3.5 FP8 on H200 at $0.218400 per million completion tokens in Experiment 006 to Qwen 3.6 source NVFP4 at $0.125959 in Experiment 332](https://nosible.com/images/2026/08/qwen-35b-a3b-serving-campaign-hero.svg) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) ## [What 350+ Experiments Taught Us About Serving Qwen 3.6 35B-A3B](https://nosible.com/blog/what-350-experiments-taught-us-about-serving-qwen-3-6) We ran more than 350 35B-A3B serving experiments over five days across Qwen 3.5 and Qwen 3.6. The final Qwen 3.6 source-NVFP4 branch cut steady-state cost from $0.218400/M to $0.125959/M. 2026-08-17 15 min read ![Thirty semantic factors rendered as daily time series on a dark grid, from geopolitical risk to firm-level political risk.](https://nosible.com/images/2026/08/semantic-factors-hero.png) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ## [The Future That Could Have Been: Turning Web Text into Semantic Stock Betas](https://nosible.com/blog/the-future-that-could-have-been-turning-web-text-into-semantic-stock-betas) Turn web text into transparent daily risk factors and stock-specific betas with an open-source, reproducible Python workflow. 2026-08-05 11 min read ![Signed contrast number line separating systemic and idiosyncratic risk](https://nosible.com/images/2026/06/walkthrough_step9_number_line.png) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ## [Two Tricks for Turning Sentence Embeddings into Clean Features](https://nosible.com/blog/the-contrastive-geometry-of-risk) A ten-step, training-free walkthrough that turns a frozen OpenAI text embedding into clean classifications: a multiclass relevance score sorts events into local, national, and global buckets, and a contrastive binary score splits systemic from idiosyncratic risk. Verified on real warnings from NOSIBLE World, the geometry matches Google's gemini-2.5-flash while staying deterministic, auditable, and effectively free. 2026-06-18 14 min read ![Daily NOSIBLE Trade Policy Uncertainty index compared with the published index](https://nosible.com/images/2026/06/nosible-tpu-daily.png) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ## [An Embedding-Based Approach to Trade and Economic Policy Uncertainty](https://nosible.com/blog/an-embedding-based-approach-to-trade-and-economic-policy-uncertainty) The Fed's Trade Policy Uncertainty index counts keywords across seven newspapers. We rebuilt it from 14.9 million NOSIBLE World events using only embeddings and five sentences, no keywords. It matches the published benchmark at 0.87 on monthly levels and 0.82 on monthly changes, as closely as the two official versions match each other. The same method, extended to sixty sentences, rebuilds the broader Economic Policy Uncertainty index and its national-security and healthcare categories. 2026-06-17 23 min read ![Running sprinter illustration representing efficient financial-sentiment model training](https://nosible.com/blog/illustrations/the-sprinter.png) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [Financial Sentiment](https://nosible.com/blog/tag/financial-sentiment) ## [Matching GPT-5.1 at Financial Sentiment with Active Learning and Qwen3](https://nosible.com/blog/fast-enough-to-matter-productionizing-tiny-transformers-for-signal-extraction) Here's how we fine-tuned Qwen3 0.6B to beat FinBERT and match GPT-5.1 accuracy. Complete with open-source models, datasets, and training scripts. Spoiler alert: active learning is all you need. 2025-12-12 27 min read ![Cybernaut-1 illustration representing agentic search with Monte Carlo Tree Search](https://nosible.com/blog/illustrations/cyber.png) [Web Search](https://nosible.com/blog/tag/web-search) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) ## [Introducing Cybernaut-1: Agentic Search using MCTS](https://nosible.com/blog/introducing-cybernaut-1-agentic-search-with-mcts) Cybernaut-1 combines our powerful hybrid-3 search algorithm with LLM-guided Monte Carlo Tree Search to deliver world class search results on difficult queries. 2025-08-26 2 min read ![Railway-track illustration representing the road to Cybernaut-1](https://nosible.com/blog/illustrations/track.png) [Web Search](https://nosible.com/blog/tag/web-search) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) ## [The Road to Cybernaut-1: Rebuilding Search for AI](https://nosible.com/blog/the-road-to-cybernaut-1) AI needs its own search engine. This is how we’re rebuilding search for AI -- and the road to Cybernaut-1, the first high-trust agentic search engine. 2025-08-20 17 min read ![3D cube illustration representing LLM ensemble distillation](https://nosible.com/blog/illustrations/cube.png) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [Financial Sentiment](https://nosible.com/blog/tag/financial-sentiment) ## [A Pattern for Scaling the Value Proposition of LLMs: Ensemble and Distil 🚀](https://nosible.com/blog/ensemble-and-distil) We introduce the ensemble and distil data pattern and use it to fit an ordinary least squares linear regression that outperforms GPT-4 at financial news sentiment classification using sentence transformer embeddings as features. 2024-02-06 12 min read [All Research](https://nosible.com/blog) > Research on AI systems, agents, and language-model applications for search and quantitative research. **URL:** https://nosible.com/blog/tag/artificial-intelligence --- --- title: "Machine Learning" description: "Research on machine-learning methods, embeddings, and model development for web and market intelligence." url: "https://nosible.com/blog/tag/machine-learning" --- /res Research # Machine Learning Research on machine-learning methods, embeddings, and model development for web and market intelligence. [All](https://nosible.com/blog) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Financial Sentiment](https://nosible.com/blog/tag/financial-sentiment) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [Trading Signals](https://nosible.com/blog/tag/trading-signals) [Web Search](https://nosible.com/blog/tag/web-search) 8 articles ![High-resolution chart of selected 35B-A3B serving records across a 350-plus experiment campaign, from Qwen 3.5 FP8 on H200 at $0.218400 per million completion tokens in Experiment 006 to Qwen 3.6 source NVFP4 at $0.125959 in Experiment 332](https://nosible.com/images/2026/08/qwen-35b-a3b-serving-campaign-hero.svg) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) ## [What 350+ Experiments Taught Us About Serving Qwen 3.6 35B-A3B](https://nosible.com/blog/what-350-experiments-taught-us-about-serving-qwen-3-6) We ran more than 350 35B-A3B serving experiments over five days across Qwen 3.5 and Qwen 3.6. The final Qwen 3.6 source-NVFP4 branch cut steady-state cost from $0.218400/M to $0.125959/M. 2026-08-17 15 min read ![Thirty semantic factors rendered as daily time series on a dark grid, from geopolitical risk to firm-level political risk.](https://nosible.com/images/2026/08/semantic-factors-hero.png) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ## [The Future That Could Have Been: Turning Web Text into Semantic Stock Betas](https://nosible.com/blog/the-future-that-could-have-been-turning-web-text-into-semantic-stock-betas) Turn web text into transparent daily risk factors and stock-specific betas with an open-source, reproducible Python workflow. 2026-08-05 11 min read ![Signed contrast number line separating systemic and idiosyncratic risk](https://nosible.com/images/2026/06/walkthrough_step9_number_line.png) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ## [Two Tricks for Turning Sentence Embeddings into Clean Features](https://nosible.com/blog/the-contrastive-geometry-of-risk) A ten-step, training-free walkthrough that turns a frozen OpenAI text embedding into clean classifications: a multiclass relevance score sorts events into local, national, and global buckets, and a contrastive binary score splits systemic from idiosyncratic risk. Verified on real warnings from NOSIBLE World, the geometry matches Google's gemini-2.5-flash while staying deterministic, auditable, and effectively free. 2026-06-18 14 min read ![Daily NOSIBLE Trade Policy Uncertainty index compared with the published index](https://nosible.com/images/2026/06/nosible-tpu-daily.png) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ## [An Embedding-Based Approach to Trade and Economic Policy Uncertainty](https://nosible.com/blog/an-embedding-based-approach-to-trade-and-economic-policy-uncertainty) The Fed's Trade Policy Uncertainty index counts keywords across seven newspapers. We rebuilt it from 14.9 million NOSIBLE World events using only embeddings and five sentences, no keywords. It matches the published benchmark at 0.87 on monthly levels and 0.82 on monthly changes, as closely as the two official versions match each other. The same method, extended to sixty sentences, rebuilds the broader Economic Policy Uncertainty index and its national-security and healthcare categories. 2026-06-17 23 min read ![Running sprinter illustration representing efficient financial-sentiment model training](https://nosible.com/blog/illustrations/the-sprinter.png) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [Financial Sentiment](https://nosible.com/blog/tag/financial-sentiment) ## [Matching GPT-5.1 at Financial Sentiment with Active Learning and Qwen3](https://nosible.com/blog/fast-enough-to-matter-productionizing-tiny-transformers-for-signal-extraction) Here's how we fine-tuned Qwen3 0.6B to beat FinBERT and match GPT-5.1 accuracy. Complete with open-source models, datasets, and training scripts. Spoiler alert: active learning is all you need. 2025-12-12 27 min read ![Cybernaut-1 illustration representing agentic search with Monte Carlo Tree Search](https://nosible.com/blog/illustrations/cyber.png) [Web Search](https://nosible.com/blog/tag/web-search) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) ## [Introducing Cybernaut-1: Agentic Search using MCTS](https://nosible.com/blog/introducing-cybernaut-1-agentic-search-with-mcts) Cybernaut-1 combines our powerful hybrid-3 search algorithm with LLM-guided Monte Carlo Tree Search to deliver world class search results on difficult queries. 2025-08-26 2 min read ![Railway-track illustration representing the road to Cybernaut-1](https://nosible.com/blog/illustrations/track.png) [Web Search](https://nosible.com/blog/tag/web-search) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) ## [The Road to Cybernaut-1: Rebuilding Search for AI](https://nosible.com/blog/the-road-to-cybernaut-1) AI needs its own search engine. This is how we’re rebuilding search for AI -- and the road to Cybernaut-1, the first high-trust agentic search engine. 2025-08-20 17 min read ![3D cube illustration representing LLM ensemble distillation](https://nosible.com/blog/illustrations/cube.png) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [Financial Sentiment](https://nosible.com/blog/tag/financial-sentiment) ## [A Pattern for Scaling the Value Proposition of LLMs: Ensemble and Distil 🚀](https://nosible.com/blog/ensemble-and-distil) We introduce the ensemble and distil data pattern and use it to fit an ordinary least squares linear regression that outperforms GPT-4 at financial news sentiment classification using sentence transformer embeddings as features. 2024-02-06 12 min read [All Research](https://nosible.com/blog) > Research on machine-learning methods, embeddings, and model development for web and market intelligence. **URL:** https://nosible.com/blog/tag/machine-learning --- --- title: "What 350+ Experiments Taught Us About Serving Qwen 3.6 35B-A3B" description: "350+ Qwen 35B-A3B serving experiments on H200 and RTX PRO 6000 cut steady-state cost 42.3%. Results cover quantization, MTP, DFlash, memory, and kernels." url: "https://nosible.com/blog/what-350-experiments-taught-us-about-serving-qwen-3-6" --- [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) # What 350+ Experiments Taught Us About Serving Qwen 3.6 35B-A3B [Stuart Reid](https://www.linkedin.com/in/stuartgordonreid/) 2026-08-17 15 min read Copy as Markdown ![High-resolution chart of selected 35B-A3B serving records across a 350-plus experiment campaign, from Qwen 3.5 FP8 on H200 at $0.218400 per million completion tokens in Experiment 006 to Qwen 3.6 source NVFP4 at $0.125959 in Experiment 332](https://nosible.com/images/2026/08/qwen-35b-a3b-serving-campaign-hero.svg) NOSIBLE World is a web-scale, point-in-time search and event system. It retrieves dated open-web sources and produces structured records with events, entities, and embeddings. World V1 processed hundreds of billions of text tokens. World V2 will cover a longer history, include more events from categories intentionally dampened in V1 to control volume, and process two to three times as much text. That workload makes the cost of Qwen serving a V2 launch constraint. We chose the open [Qwen 35B-A3B](https://huggingface.co/Qwen/Qwen3.6-35B-A3B-FP8) model because no model we assessed delivered higher task quality at lower serving cost. Earlier this year, we benchmarked it against [Gemini Flash](https://ai.google.dev/gemini-api/docs/models) and [DeepSeek](https://github.com/deepseek-ai) for creating NOSIBLE World. Qwen matched or exceeded both on that workload. The benchmark selected the Qwen model family; this campaign selected its serving configuration. We hope Qwen releases a 3.8 35B-A3B variant soon because it would be the first model we test with this serving method for World V2. Five days and more than 350 serving experiments reduced the steady-state cost of the Qwen 35B-A3B serving branch used to build World from $0.218400/M to $0.125959/M. The final branch generated 6,064.60 completion tokens per second on one [RTX PRO 6000 Blackwell](https://www.nvidia.com/en-us/products/workstations/professional-desktop-gpus/rtx-pro-6000/), using a $2.75-per-hour GPU-cost normalisation. The campaign covered [Qwen 3.5](https://huggingface.co/Qwen/Qwen3.5-35B-A3B-FP8) and [Qwen 3.6](https://huggingface.co/Qwen/Qwen3.6-35B-A3B-FP8), and every result retains its actual model label. ## 1. Checkpoint packages and quantisation Qwen 3.6 35B-A3B is a mixture-of-experts model with 35 billion total parameters and roughly three billion active parameters per generated token. The checkpoint's routed-expert layout and the kernel that executes it therefore affect serving cost. We screened weight packages on H200 and RTX PRO 6000 to choose branches for deeper work. The results below use their recorded serving conditions and do not rank every checkpoint for every workload. | Checkpoint family | Representative result | What it told us | | --- | --- | --- | | [Intel mixed-INT4 AutoRound](https://huggingface.co/Intel/Qwen3.6-35B-A3B-int4-mixed-AutoRound) | 5,126.77 tok/s; $0.149000/M | The initial production benchmark to beat. | | [Unsloth NVFP4 Fast](https://huggingface.co/unsloth/Qwen3.6-35B-A3B-NVFP4-Fast) | 5,040.22 tok/s; $0.151559/M | Four-bit floating point did not win without tuning. | | [Red Hat AI NVFP4](https://huggingface.co/RedHatAI/Qwen3.6-35B-A3B-NVFP4) | 4,979.54 tok/s; $0.153405/M | A valid alternative, but not the first priority. | | [Red Hat AI dynamic FP8](https://huggingface.co/RedHatAI/Qwen3.6-35B-A3B-FP8-dynamic) | 7,019.22 tok/s; $0.197870/M on H200 | Valid H200 package screen; it did not advance. | | [QuantTrio AWQ](https://huggingface.co/QuantTrio/Qwen3.6-35B-A3B-AWQ) | 4,957.00 tok/s; $0.154103/M | Usable, but behind the early leaders. | | [RDTand PrismaQuant](https://huggingface.co/rdtand/Qwen3.6-35B-A3B-PrismaQuant-4.75bit-vllm) | 4,616.41 tok/s; $0.165473/M | Did not justify more work in this campaign. | | [88plug W4A16](https://huggingface.co/88plug/Qwen3.6-35B-A3B-W4A16) | 4,435.00 tok/s; $0.172241/M | Lower precision alone did not overcome the serving cost. | NVFP4 uses four-bit floating-point weights. W4A16 uses four-bit weights and 16-bit activations. Those labels omit the execution library, routed-expert layout, auxiliary modules, and available memory that determined these results. A weight-format label does not predict serving cost. On [H200](https://www.nvidia.com/en-us/data-center/h200/), [NVIDIA's ModelOpt NVFP4 Qwen checkpoint](https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4) reached 6,623.18 completion tokens per second at $0.209701/M; the same path on RTX PRO 6000 reached 2,563.40 tokens per second at $0.297998/M. On the RTX PRO 6000 run, the [Marlin](https://github.com/IST-DASLab/marlin) path reported no native FP4 compute and used weight-only FP4 compression. Reducing native MTP depth from three proposals to one on H200 improved the package to 7,022.95 tokens per second and $0.197764/M. The [FlashInfer](https://github.com/flashinfer-ai/flashinfer) TensorRT-LLM mixture-of-experts alternative failed at engine startup on the deployed Blackwell configuration. ### Intel's mixed-INT4 AutoRound was the benchmark [Intel's Qwen3.6-35B-A3B-int4-mixed-AutoRound checkpoint](https://huggingface.co/Intel/Qwen3.6-35B-A3B-int4-mixed-AutoRound) kept 243 precision-sensitive modules in FP16 and quantised routed experts to four-bit weights. It loaded, completed the offered client load, and executed on [Marlin](https://github.com/IST-DASLab/marlin). It set the Qwen 3.6 cost benchmark during the checkpoint search. | Change from the early Intel configuration | Result | Decision | | --- | --- | --- | | Reduce resident sequences from 1,280 to 512 | $0.149726/M | Better scheduling geometry | | Reduce batch capacity from 24,576 to 16,384 tokens | $0.149179/M | Better prefill/decode balance | | Use the CUDA 12.9 vLLM build | $0.149000/M | Canonical Intel configuration | | Force the Marlin mixture-of-experts backend | $0.148947/M | 0.036% numerical edge; not adopted | Resident sequences count requests scheduled on the GPU; they exclude clients waiting for work. The explicit-Marlin A/B measured 5,128.61 tokens per second and $0.148947/M, 0.036% below the automatic CUDA 12.9 control. It missed the 0.75% adoption threshold. Later Intel sweeps produced lower isolated values, including $0.148162/M, but did not replace the replicated automatic control. Experiment 142 remained the canonical Intel configuration. [AutoRound](https://github.com/intel/auto-round) documents the quantisation method; the checkpoint card lists the retained modules. Intel's mixed-INT2 branch required a runtime repair because Qwen's auxiliary predictor contains BF16 tensors while routed experts use the low-bit Humming path. After that repair, two full passes reached 5,218.81 tokens per second and $0.146372/M. The branch did not become the final economic leader. Screen a quantised checkpoint as a package: retained precision, kernel path, auxiliary-module compatibility, and sustained-load behaviour determine the result alongside bit width. ## 2. Hardware cost The campaign began on [H200](https://www.nvidia.com/en-us/data-center/h200/) and later screened [RTX PRO 6000 Blackwell](https://www.nvidia.com/en-us/products/workstations/professional-desktop-gpus/rtx-pro-6000/). Experiment 006 is the Qwen 3.5 campaign starting point. It is outside the Qwen 3.6 GPU comparison. H200 produced more raw completion tokens per second in these records, while RTX PRO 6000 produced the lower cost per completion token. | Representative branch | GPU | Completion tok/s | Cost basis used in the record | Decision | | --- | --- | --- | --- | --- | | Qwen 3.5 FP8, Experiment 006 | H200 | 6,359.80 | $0.218400/M at recorded $5/h | Historical campaign starting point | | Qwen 3.6 Intel mixed-INT4, Experiment 087 | H200 | 8,061.94 | $0.172277/M at recorded $5/h | Initial Qwen 3.6 Intel H200 screen | | Qwen 3.6 [PalmFuture GPTQ](https://huggingface.co/palmfuture/Qwen3.6-35B-A3B-GPTQ-Int4) | H200 | 8,489.40 | $0.163603/M at recorded $5/h | Best H200 screen in this table | | Qwen 3.6 Intel mixed-INT4 | H200 | 8,377.97 | $0.165779/M at recorded $5/h | Strong throughput, not the cost leader | | Qwen 3.6 Intel mixed-INT4 | RTX PRO 6000 | 5,126.77 | $0.149000/M at $2.75/h normalisation | Qwen 3.6 leader at the time | | Qwen 3.6 source NVFP4 | RTX PRO 6000 | 6,064.60 | $0.125959/M at $2.75/h normalisation | Final steady-state performance leader | The H200 and RTX records use different serving conditions and cost bases. They are screening outcomes, not a matched GPU A/B. Calculate cost per completion token with the hourly rate you expect to pay before buying a faster GPU. Include latency targets, replica count, and utilisation when those drive the deployment decision. ## 3. KV cache and memory capacity The key-value cache stores context needed to generate the next token. Cache format and available GPU memory therefore determine how much request state the service can retain. In an otherwise identical Intel MTP1/Marlin configuration, [FP8 E4M3 key-value cache](https://docs.vllm.ai/en/v0.5.0/quantization/fp8_e4m3_kvcache.html) raised throughput from 4,940.76 to 5,072.34 tokens per second and lowered cost from $0.154610/M to $0.150599/M. Changing its default 16-token pages to 32 tokens reduced throughput to 5,018.22 tokens per second and raised cost to $0.152223/M. For managed NVFP4, we held the checkpoint and final replay fixture fixed and changed GPU memory available to the runtime. | GPU-memory utilisation | Completion tok/s | $/M completion tokens | Decision | | --- | --- | --- | --- | | 0.92 | 5,564.91 | $0.137269 | Control | | 0.94 | 5,519.31 | $0.138403 | Rejected: worse than 0.92 | | 0.96 | **5,684.28** | **$0.134386** | Adopted | The 0.96 configuration was 2.10% cheaper than the 0.92 control and 9.81% cheaper than Intel's canonical $0.149000/M result. Moving from 0.92 to 0.94 raised $/M; moving from 0.94 to 0.96 lowered $/M. The final path used FP8 E4M3 cache, an eight-bit float with four exponent bits and three mantissa bits. ## 4. Admission control and scheduler policy Admission and scheduling controls set the operating limits: - Reducing resident capacity from 512 to 480 lowered $/M by 0.48%. The gain missed the adoption threshold, so 512 remained the operating setting. - At 1,536 offered clients, $/M rose by 0.91% and mean time to first token increased by 3.76 seconds. At 1,024 offered clients, $/M rose by 0.44% and mean time to first token fell by 3.81 seconds. The 1,280-client setting remained the cost choice for this batch service. - Reducing the model-length ceiling from 2,048 to 1,024 tokens produced 5,110.34 tokens per second and $0.149479/M in the Intel control. It was 0.32% slower and more expensive, and it excluded World inputs that span 1,024 to 2,048 tokens. - Batch-token limits of 12,288 and 20,480 cost $0.135746/M and $0.135406/M, respectively, versus $0.134386/M at 16,384. - A 24,576-token setting failed the warm-up gate with a CUDA out-of-memory error before a full score could exist. - A dense [CUDA Graph](https://docs.nvidia.com/cuda/cuda-c-programming-guide/#cuda-graphs) ladder shaped around observed traffic was 3.42% slower than the standard ladder. - Disabling prefix caching, chunked prefill, or statistics failed to clear the 0.75% adoption threshold. Requests for alternate asynchronous scheduling produced no material gain, but logs did not confirm that the alternate scheduler activated. Cache format, memory fraction, resident capacity, batch capacity, prefill policy, and graph capture compete for GPU memory and scheduling time. Longer prompts increase cache pressure. Longer completions change the prefill/decode mix. Latency-sensitive services can prefer a capacity setting that a saturated batch service rejects. Select the cheapest setting that completes the required workload. ## 5. Native MTP and NTP Standard next-token prediction (NTP) generates one token per decoding step. Qwen's native multi-token prediction (MTP) path adds a predictor that proposes tokens for the full model to accept or reject. We tested the NTP baseline, MTP with one, two, and three proposals, and a language-model-only MTP control. Accepted proposals must save more target-model work than the predictor adds in compute and memory overhead. | Matched configuration | Completion tok/s | $/M completion tokens | Outcome | | --- | --- | --- | --- | | Managed NVFP4 with one proposal | 5,658.41 | $0.135001 | Control | | Same configuration without the predictor | 5,033.49 | $0.151761 | 11.04% lower throughput | | Same configuration, one proposal, text-only mode | 5,092.55 | $0.150001 | 10.00% lower throughput | Removing the MTP predictor (the NTP baseline) reduced matched-control throughput by 11.04%. In the earlier configuration with 0.92 GPU-memory utilisation and a 24,576-token batch limit, one MTP proposal cost $0.151559/M and two proposals cost $0.160839/M. The second proposal raised cost. In an otherwise matched Intel test, three proposals reached 4,534.37 tokens per second: 6.91% slower and 7.43% more expensive than one proposal. ## 6. DFlash: not viable for World DFlash is a separate [block-diffusion speculative-decoding method](https://github.com/z-lab/dflash). It uses the official FP8 target and an external eight-token draft. It does not use Qwen's native MTP head. We tested it as a separate serving mode on [SGLang](https://github.com/sgl-project/sglang). | Attempt | Outcome | What it established | | --- | --- | --- | | H200, 1,280-resident target path | No traffic: target attention required SM100. | The target path was not H200-compatible. | | H200, FA3 fallback at 1,280 resident requests | No traffic: empty startup exit. | The failure had no defensible single cause. | | H200, minimal 700-resident path | No traffic: platform diagnostics reported OOM, but raw logs did not confirm it. | This is a capacity diagnostic, not a throughput result. | | H200, published 32-resident capacity shape | 3,583.78 tok/s; $0.387549/M; 4,096 / 4,096 requests completed. | It was valid only after capacity fell to 32 resident requests. | | RTX PRO 6000, published target path | No traffic: target attention required SM100. | The path was not compatible with this RTX Blackwell GPU. | | RTX PRO 6000, FA4 fallback | No traffic: startup required BF16 Mamba state. | The failure specified a repair. | | RTX PRO 6000, FA4 fallback with the repair | No traffic: empty startup exit. | The method produced no RTX performance result. | The valid H200 run queued 1,280 offered clients behind 32 resident requests. It delivered a real full-load result, but at $0.387549/M it was not cost-competitive for World. No no-draft control used the same SGLang runtime and 32-resident shape, so this result rejects the deployment configuration; it does not isolate DFlash overhead. The high-capacity H200 route and every RTX route stopped at compatibility or startup gates. DFlash did not enter the deployment choice. Draft acceptance, draft-state memory, target-GPU kernel support, and scheduler capacity can each eliminate a net speculation gain. Compare each separate draft method with a matched no-speculation control. ## 7. Execution paths A kernel is the GPU program that performs a model operation. [Marlin](https://github.com/IST-DASLab/marlin), [FlashInfer](https://github.com/flashinfer-ai/flashinfer), [CUTLASS](https://github.com/NVIDIA/cutlass), [Triton](https://github.com/triton-lang/triton), [TensorRT-LLM](https://github.com/NVIDIA/TensorRT-LLM), and [DeepGEMM](https://github.com/deepseek-ai/DeepGEMM) provide low-precision execution paths. A runtime can select a path, fall back to another, or reject a layout. In a controlled [official Qwen 3.5 FP8](https://huggingface.co/Qwen/Qwen3.5-35B-A3B-FP8) H200 test, enabling DeepGEMM delivered 7,197.30 completion tokens per second and $0.192974/M. Disabling only DeepGEMM delivered 7,661.29 tokens per second and $0.181286/M: 6.45% more throughput and 6.06% lower cost. The kernel remained compatible; it was slower for this workload. ### B12X mixture-of-experts routing B12X was a separate [FlashInfer](https://github.com/flashinfer-ai/flashinfer) execution route for the NVFP4 target's mixture-of-experts layers. It selected successfully and completed World traffic, but it was expensive in every early configuration. | B12X configuration | $/M completion tokens | Outcome | | --- | --- | --- | | Target B12X with MTP3 | $0.282431 | Completed, but uneconomic | | Target B12X with MTP2 | $0.254162 | Completed, but uneconomic | | Target B12X with MTP2 and Marlin atomic add | $0.258342 | Completed, but more expensive | | Target B12X with MTP1 | $0.234710 | Best B12X result; still uneconomic | Later source tests could select B12X but could not carry the full World load. One stopped before API readiness, two could not allocate a 2.58 GiB minimum cache, a reduced-batch warm-up timed out 168 of 1,024 clients, and the first scored pass completed 1,520 of 4,096 requests. B12X did not advance. Unsloth NVFP4 began at 5,040.22 tokens per second and $0.151559/M, behind Intel. Memory and scheduler changes lifted managed NVFP4 to 5,684.28 tokens per second and $0.134386/M. We ran the source runtime in fresh allocations and averaged two passes at 6,064.60 tokens per second and $0.125959/M with the same checkpoint. The final runtime used [vLLM](https://github.com/vllm-project/vllm) source commit 226749b, built with CUDA 13 for Blackwell. Startup logs confirmed the selected path: - FlashInfer CUTLASS for NVFP4 linear operations. - FlashInfer CUTLASS for NVFP4 mixture-of-experts operations. - FlashInfer attention. - FlashInfer CUTLASS for the unquantised auxiliary predictor. The source runtime used [Triton](https://github.com/triton-lang/triton) kernels v3.5.1 and a Blackwell adaptation that disabled an unsupported [FlashAttention-3](https://github.com/Dao-AILab/flash-attention) (FA3) extension. The adaptation changes the build and selected attention path. Runtime logs selected FlashInfer attention. Alternate execution paths either rejected the NVFP4 checkpoint, failed the 1,280-client load, or cost more. An exact-shape mixture-of-experts autotuning test changed the maximum graph-capture size for FlashInfer CUTLASS autotuning. It completed at 5,990.35 tokens per second and $0.127520/M, 1.24% more expensive than the automatic final path. The source winner completed two full fresh-allocation passes. | Source-runtime pass | Successful requests | Completion tok/s | $/M completion tokens | | --- | --- | --- | --- | | Pass 1 | 4,096 / 4,096 | 6,035.68 | $0.126562 | | Pass 2 | 4,096 / 4,096 | 6,093.53 | $0.125359 | | Mean | 8,192 / 8,192 | **6,064.60** | **$0.125959** | This result was 6.27% cheaper than managed NVFP4 and 15.46% cheaper than Intel's canonical configuration. We selected the automatic backend after reviewing startup logs and reproducing the result in a fresh allocation. ## 8. Transport and release gates Delivery mode and client concurrency can change a serving result under offered load. No delivery variant produced a repeatable performance gain, so none entered the final configuration. Managed NVFP4 completed an endurance run of 16,384 requests without HTTP or streaming errors. Its deterministic translation screen covered 16 source languages and found no blank output, repetition loop, malformed response, or observed systematic omission of named entities or numbers. The source runtime has throughput evidence, not output qualification. It must pass the same output screen before promotion. ### Release decisions | Question | Required evidence | | --- | --- | | Can it load? | Startup and functional response | | Can it carry intended demand? | High-concurrency warm-up and full request completion | | Is it cheaper? | Repeated, matched steady-state measurement | | Is a gain real? | Fresh-allocation reproduction | | Is it safe to release? | Endurance and representative output checks | ## Test protocol and adoption rules In the matched final RTX fixture, [Qwen 3.5 Intel AutoRound](https://huggingface.co/Intel/Qwen3.5-35B-A3B-int4-AutoRound) cost $0.149374/M and [Qwen 3.6 Intel AutoRound](https://huggingface.co/Intel/Qwen3.6-35B-A3B-int4-mixed-AutoRound) cost $0.149000/M. The 0.25% difference missed the 0.75% adoption threshold, so we treat the releases as one serving family for this non-thinking workload. Each experiment answered a bounded question: whether a checkpoint loaded, whether a kernel supported its layout, whether a serving shape survived warm-up, or whether every client request completed. A failed gate ended that branch before it consumed full-load test time. The final Blackwell comparison records replayed NOSIBLE World translation requests with a 2,048-token context limit, streaming output, disabled thinking, a 256-token completion limit, 1,280 offered clients, and two measured 4,096-request passes. Every request had to finish. Earlier feasibility, package, and H200 screens chose candidates. They were not matched final comparisons. The final result applies to this replay fixture. Different GPU pricing, prompt lengths, completion lengths, offered concurrency, latency limits, or quality requirements require rerunning the relevant tests. The campaign used five gates: 1. A functional request confirmed a valid response. 2. A high-concurrency warm-up exposed capacity and stability problems. 3. Two full passes established throughput and cost before a late-stage candidate could change the recommendation. 4. A fresh allocation reproduced material gains. 5. Endurance and output checks qualified a release candidate. The headline reports normalised steady-state GPU cost: ``` normalised cost per million completion tokens = hourly GPU cost x 1,000,000 / (completion tokens per second x 3,600) ``` The final RTX comparison uses $2.75 per hour as a campaign normalisation. Experiment 332 used an allocation priced at $2.09 per GPU hour plus $0.021 per disk hour, so $0.125959/M is not a provider-neutral deployment quote. H200 records retain their recorded $5-per-hour rate. The calculation excludes startup, source compilation, model download, and JIT work; short-lived jobs must amortise those costs separately. We adopted only replicated improvements of at least 0.75%. An apparent gain above 3% required a fresh-allocation reproduction. Smaller measured differences stayed in the record but did not change the configuration. ![Measurement gates for the Qwen serving campaign](https://nosible.com/images/2026/08/qwen-measurement-gates.svg) Measurement gates for the Qwen serving campaign ## The decision map ![Decision ledger for the major Qwen serving choices](https://nosible.com/images/2026/08/qwen-decision-ledger.svg) Decision ledger for the major Qwen serving choices | Family | Work conducted | Final status | | --- | --- | --- | | Runtime compatibility | Managed images, newer runtime images, source builds, startup and model-load gates | Only clean services advanced to throughput work | | Checkpoints and weight formats | [Official FP8](https://huggingface.co/Qwen/Qwen3.6-35B-A3B-FP8), [GPTQ](https://huggingface.co/palmfuture/Qwen3.6-35B-A3B-GPTQ-Int4), [AWQ](https://huggingface.co/QuantTrio/Qwen3.6-35B-A3B-AWQ), [NVIDIA NVFP4](https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4), [Unsloth NVFP4](https://huggingface.co/unsloth/Qwen3.6-35B-A3B-NVFP4-Fast), [W4A16](https://huggingface.co/88plug/Qwen3.6-35B-A3B-W4A16), and [Intel mixed-INT4](https://huggingface.co/Intel/Qwen3.6-35B-A3B-int4-mixed-AutoRound) packages | NVFP4 source performance path selected; Intel mixed-INT4 retained as the benchmark; official NVIDIA NVFP4 did not advance; source output gate pending | | GPU and provider economics | [H200](https://www.nvidia.com/en-us/data-center/h200/) and [RTX PRO 6000](https://www.nvidia.com/en-us/products/workstations/professional-desktop-gpus/rtx-pro-6000/) screening using recorded H200 rates and the RTX campaign normalization | RTX PRO 6000 won the operating-cost decision | | Quantised execution | [Marlin](https://github.com/IST-DASLab/marlin), Humming, [FlashInfer CUTLASS](https://github.com/flashinfer-ai/flashinfer), [Triton](https://github.com/triton-lang/triton), [DeepGEMM](https://github.com/deepseek-ai/DeepGEMM), and other candidate paths | Use the selected path confirmed in startup logs; DeepGEMM disabled in the tested H200 path | | Context state and cache | [FP8 key-value cache](https://docs.vllm.ai/en/v0.5.0/quantization/fp8_e4m3_kvcache.html), 16- and 32-token pages, state capacity, cache format and cache allocation | FP8 E4M3 retained on the final path; 16-token pages retained in the Intel MTP1/Marlin control | | Context, memory and admission | Model-length ceiling, GPU-memory fraction, resident sequences, offered concurrency, batch-token ceiling, time to first token | 2,048 context, 0.96 memory, 512 resident sequences, 16,384 batch tokens, C1280 cost choice | | Scheduler and graph policy | [Chunked prefill and prefix cache](https://docs.vllm.ai/en/latest/features/automatic_prefix_caching.html), asynchronous modes, [CUDA Graph](https://docs.nvidia.com/cuda/cuda-c-programming-guide/#cuda-graphs) coverage | Keep the validated standard path; dense graph override rejected | | Native MTP and NTP | [No proposal, one proposal, two proposals, three proposals, and text-only control](https://docs.vllm.ai/en/latest/features/spec_decode/) | MTP1 adopted; MTP2 and MTP3 rejected | | DFlash speculation | [Official FP8 target with an eight-token draft](https://huggingface.co/z-lab/Qwen3.6-35B-A3B-DFlash); compatibility and capacity variants | Valid full-load result, but not competitive | | Request delivery | Streaming and transport variants under offered client load | No adopted improvement | | Reproducibility | Repeated passes, fresh allocations, backend selection logs and source-build audit | Required for the final source-runtime result | | Qualification | Functional, warm-up, full-load, endurance and output checks | Source output screen remains the final release gate | ## Actions for another serving campaign 1. Choose GPUs based on output cost, latency, capacity, and quality requirements. 2. Screen checkpoint packages before deep parameter sweeps; a mixed-precision checkpoint can beat a lower-bit format. 3. Tune cache format, memory fraction, batch limit, resident capacity, and graph capture together because they use the same GPU memory and scheduling resources. 4. Compare native MTP with its NTP control, then test external draft methods as separate serving modes. 5. Record the runtime path and build adaptations because kernel selection can change measured throughput and $/M. For the fixed NOSIBLE World replay fixture, Unsloth NVFP4 Fast on RTX PRO 6000 Blackwell was the steady-state performance leader at 6,064.60 completion tokens per second and $0.125959/M. The source runtime used one native proposal, FP8 E4M3 cache, 0.96 GPU-memory utilisation, 16,384 batch tokens, and 512 resident sequences. The source runtime must pass the retained output screen before release. World V2 is why we ran this campaign. Re-run the relevant branches when GPU architecture, GPU price, prompt length, completion length, offered concurrency, latency target, or quality requirements change. For a new Blackwell stack, start with the [NVIDIA RTX PRO 6000 specifications](https://www.nvidia.com/en-us/products/workstations/professional-desktop-gpus/rtx-pro-6000/), the [CUDA compatibility guide](https://docs.nvidia.com/cuda/blackwell-compatibility-guide/index.html), and the [NVIDIA Developer Forums](https://forums.developer.nvidia.com/), then measure the service. ## External sources and implementation terrain - [Qwen 3.5 35B-A3B FP8 model card](https://huggingface.co/Qwen/Qwen3.5-35B-A3B-FP8) and [Qwen 3.6 35B-A3B FP8 model card](https://huggingface.co/Qwen/Qwen3.6-35B-A3B-FP8) - [NVIDIA Qwen 3.6 NVFP4 checkpoint](https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4), [Intel mixed-INT4 AutoRound checkpoint](https://huggingface.co/Intel/Qwen3.6-35B-A3B-int4-mixed-AutoRound), and [Unsloth NVFP4 Fast checkpoint](https://huggingface.co/unsloth/Qwen3.6-35B-A3B-NVFP4-Fast) - [NVIDIA Model Optimizer NVFP4 reference](https://nvidia.github.io/Model-Optimizer/reference/generated/modelopt.torch.kernels.quantization.common.nvfp4_quant.html), [AutoRound](https://github.com/intel/auto-round), and [Marlin](https://github.com/IST-DASLab/marlin) - [vLLM](https://github.com/vllm-project/vllm), [vLLM speculative decoding](https://docs.vllm.ai/en/latest/features/spec_decode/), [vLLM FP8 key-value cache](https://docs.vllm.ai/en/v0.5.0/quantization/fp8_e4m3_kvcache.html), [FlashInfer](https://github.com/flashinfer-ai/flashinfer), [CUTLASS](https://github.com/NVIDIA/cutlass), [Triton](https://github.com/triton-lang/triton), [TensorRT-LLM](https://github.com/NVIDIA/TensorRT-LLM), and [DeepGEMM](https://github.com/deepseek-ai/DeepGEMM) - [DFlash](https://github.com/z-lab/dflash), [NVIDIA H200](https://www.nvidia.com/en-us/data-center/h200/), [NVIDIA RTX PRO 6000 Blackwell](https://www.nvidia.com/en-us/products/workstations/professional-desktop-gpus/rtx-pro-6000/), [Gemini models](https://ai.google.dev/gemini-api/docs/models), [DeepSeek](https://github.com/deepseek-ai), and the [NVIDIA Developer Forums](https://forums.developer.nvidia.com/) [All Research](https://nosible.com/blog) Related Research ![Thirty semantic factors rendered as daily time series on a dark grid, from geopolitical risk to firm-level political risk.](https://nosible.com/images/2026/08/semantic-factors-hero.png) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ### [The Future That Could Have Been: Turning Web Text into Semantic Stock Betas](https://nosible.com/blog/the-future-that-could-have-been-turning-web-text-into-semantic-stock-betas) 2026-08-05 11 min read ![Signed contrast number line separating systemic and idiosyncratic risk](https://nosible.com/images/2026/06/walkthrough_step9_number_line.png) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ### [Two Tricks for Turning Sentence Embeddings into Clean Features](https://nosible.com/blog/the-contrastive-geometry-of-risk) 2026-06-18 14 min read ![Daily NOSIBLE Trade Policy Uncertainty index compared with the published index](https://nosible.com/images/2026/06/nosible-tpu-daily.png) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ### [An Embedding-Based Approach to Trade and Economic Policy Uncertainty](https://nosible.com/blog/an-embedding-based-approach-to-trade-and-economic-policy-uncertainty) 2026-06-17 23 min read > 350+ Qwen 35B-A3B serving experiments on H200 and RTX PRO 6000 cut steady-state cost 42.3%. Results cover quantization, MTP, DFlash, memory, and kernels. **URL:** https://nosible.com/blog/what-350-experiments-taught-us-about-serving-qwen-3-6 --- --- title: "Risk Indicator" description: "Research on constructing and validating point-in-time indicators of economic, geopolitical, and market risk." url: "https://nosible.com/blog/tag/risk-indicator" --- /res Research # Risk Indicator Research on constructing and validating point-in-time indicators of economic, geopolitical, and market risk. [All](https://nosible.com/blog) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Financial Sentiment](https://nosible.com/blog/tag/financial-sentiment) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [Trading Signals](https://nosible.com/blog/tag/trading-signals) [Web Search](https://nosible.com/blog/tag/web-search) 5 articles ![Thirty semantic factors rendered as daily time series on a dark grid, from geopolitical risk to firm-level political risk.](https://nosible.com/images/2026/08/semantic-factors-hero.png) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ## [The Future That Could Have Been: Turning Web Text into Semantic Stock Betas](https://nosible.com/blog/the-future-that-could-have-been-turning-web-text-into-semantic-stock-betas) Turn web text into transparent daily risk factors and stock-specific betas with an open-source, reproducible Python workflow. 2026-08-05 11 min read ![Signed contrast number line separating systemic and idiosyncratic risk](https://nosible.com/images/2026/06/walkthrough_step9_number_line.png) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ## [Two Tricks for Turning Sentence Embeddings into Clean Features](https://nosible.com/blog/the-contrastive-geometry-of-risk) A ten-step, training-free walkthrough that turns a frozen OpenAI text embedding into clean classifications: a multiclass relevance score sorts events into local, national, and global buckets, and a contrastive binary score splits systemic from idiosyncratic risk. Verified on real warnings from NOSIBLE World, the geometry matches Google's gemini-2.5-flash while staying deterministic, auditable, and effectively free. 2026-06-18 14 min read ![Daily NOSIBLE Trade Policy Uncertainty index compared with the published index](https://nosible.com/images/2026/06/nosible-tpu-daily.png) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ## [An Embedding-Based Approach to Trade and Economic Policy Uncertainty](https://nosible.com/blog/an-embedding-based-approach-to-trade-and-economic-policy-uncertainty) The Fed's Trade Policy Uncertainty index counts keywords across seven newspapers. We rebuilt it from 14.9 million NOSIBLE World events using only embeddings and five sentences, no keywords. It matches the published benchmark at 0.87 on monthly levels and 0.82 on monthly changes, as closely as the two official versions match each other. The same method, extended to sixty sentences, rebuilds the broader Economic Policy Uncertainty index and its national-security and healthcare categories. 2026-06-17 23 min read ![S&P 500 news-stress overlay with drawdown and signal thresholds](https://nosible.com/images/2026/06/news-stress-overlay-hero.png) [Trading Signals](https://nosible.com/blog/tag/trading-signals) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ## [Turning News into a Risk-On/Risk-Off Equity Signal](https://nosible.com/blog/turning-news-into-a-risk-on-risk-off-equity-signal) We built a risk-on/risk-off trading signal from the NOSIBLE event database that measures how much of the global news flow is about market-stress themes, holding equities when that reading is low and moving to T-bills when it spikes. Selected on 2010 to 2013 and tested on an untouched 2015 to 2026 window, it held the S&P 500's buy-and-hold return (+254% versus +269%) while cutting the maximum drawdown from −34% to −18% and raising the Sharpe ratio from 0.64 to 0.89. The same rule transfers unchanged to the Nasdaq and the Russell 2000. 2026-06-16 9 min read ![NOSIBLE geopolitical risk signal compared with published geopolitical risk indices](https://nosible.com/images/2026/06/nosible-gpr-vs-published.png) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ## [We Rebuilt the Geopolitical Risk Index with Nosible World](https://nosible.com/blog/rebuilding-the-geopolitical-risk-index-from-nosible-world) Markets move on geopolitics, but risk models cannot read the news. We turned 13.2 million news events into a geopolitical risk signal, matched the Federal Reserve benchmark, and broke it down by country, by country pair, and into an oil supply-risk signal. 2026-06-06 15 min read [All Research](https://nosible.com/blog) > Research on constructing and validating point-in-time indicators of economic, geopolitical, and market risk. **URL:** https://nosible.com/blog/tag/risk-indicator --- --- title: "NOSIBLE World" description: "Research using NOSIBLE World, the point-in-time event dataset for web-scale quantitative analysis." url: "https://nosible.com/blog/tag/nosible-world" --- /res Research # NOSIBLE World Research using NOSIBLE World, the point-in-time event dataset for web-scale quantitative analysis. [All](https://nosible.com/blog) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Financial Sentiment](https://nosible.com/blog/tag/financial-sentiment) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [Trading Signals](https://nosible.com/blog/tag/trading-signals) [Web Search](https://nosible.com/blog/tag/web-search) 6 articles ![Thirty semantic factors rendered as daily time series on a dark grid, from geopolitical risk to firm-level political risk.](https://nosible.com/images/2026/08/semantic-factors-hero.png) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ## [The Future That Could Have Been: Turning Web Text into Semantic Stock Betas](https://nosible.com/blog/the-future-that-could-have-been-turning-web-text-into-semantic-stock-betas) Turn web text into transparent daily risk factors and stock-specific betas with an open-source, reproducible Python workflow. 2026-08-05 11 min read ![NOSIBLE World knowledge graph showing entity connections over a decade](https://nosible.com/images/2026/07/kg-hero-decade.png) [Web Search](https://nosible.com/blog/tag/web-search) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ## [Point-in-Time Knowledge Graphs over Named Entities with NOSIBLE World](https://nosible.com/blog/point-in-time-knowledge-graphs-over-named-entities) Companies change: people leave, products die, mergers happen. This post shows how to build point-in-time knowledge graphs over any company from news alone, using NOSIBLE World, which covers 38 thousand tickers, 3.2 million organizations and 6.4 million people. Named entity recognition supplies the nodes, lift-scored co-mention supplies the links, and the document date keeps every yearly view safe to backtest. 2026-07-16 8 min read ![Signed contrast number line separating systemic and idiosyncratic risk](https://nosible.com/images/2026/06/walkthrough_step9_number_line.png) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ## [Two Tricks for Turning Sentence Embeddings into Clean Features](https://nosible.com/blog/the-contrastive-geometry-of-risk) A ten-step, training-free walkthrough that turns a frozen OpenAI text embedding into clean classifications: a multiclass relevance score sorts events into local, national, and global buckets, and a contrastive binary score splits systemic from idiosyncratic risk. Verified on real warnings from NOSIBLE World, the geometry matches Google's gemini-2.5-flash while staying deterministic, auditable, and effectively free. 2026-06-18 14 min read ![Daily NOSIBLE Trade Policy Uncertainty index compared with the published index](https://nosible.com/images/2026/06/nosible-tpu-daily.png) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ## [An Embedding-Based Approach to Trade and Economic Policy Uncertainty](https://nosible.com/blog/an-embedding-based-approach-to-trade-and-economic-policy-uncertainty) The Fed's Trade Policy Uncertainty index counts keywords across seven newspapers. We rebuilt it from 14.9 million NOSIBLE World events using only embeddings and five sentences, no keywords. It matches the published benchmark at 0.87 on monthly levels and 0.82 on monthly changes, as closely as the two official versions match each other. The same method, extended to sixty sentences, rebuilds the broader Economic Policy Uncertainty index and its national-security and healthcare categories. 2026-06-17 23 min read ![S&P 500 news-stress overlay with drawdown and signal thresholds](https://nosible.com/images/2026/06/news-stress-overlay-hero.png) [Trading Signals](https://nosible.com/blog/tag/trading-signals) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ## [Turning News into a Risk-On/Risk-Off Equity Signal](https://nosible.com/blog/turning-news-into-a-risk-on-risk-off-equity-signal) We built a risk-on/risk-off trading signal from the NOSIBLE event database that measures how much of the global news flow is about market-stress themes, holding equities when that reading is low and moving to T-bills when it spikes. Selected on 2010 to 2013 and tested on an untouched 2015 to 2026 window, it held the S&P 500's buy-and-hold return (+254% versus +269%) while cutting the maximum drawdown from −34% to −18% and raising the Sharpe ratio from 0.64 to 0.89. The same rule transfers unchanged to the Nasdaq and the Russell 2000. 2026-06-16 9 min read ![NOSIBLE geopolitical risk signal compared with published geopolitical risk indices](https://nosible.com/images/2026/06/nosible-gpr-vs-published.png) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ## [We Rebuilt the Geopolitical Risk Index with Nosible World](https://nosible.com/blog/rebuilding-the-geopolitical-risk-index-from-nosible-world) Markets move on geopolitics, but risk models cannot read the news. We turned 13.2 million news events into a geopolitical risk signal, matched the Federal Reserve benchmark, and broke it down by country, by country pair, and into an oil supply-risk signal. 2026-06-06 15 min read [All Research](https://nosible.com/blog) > Research using NOSIBLE World, the point-in-time event dataset for web-scale quantitative analysis. **URL:** https://nosible.com/blog/tag/nosible-world --- --- title: "The Future That Could Have Been: Turning Web Text into Semantic Stock Betas" description: "Learn how semantic-factors turns NOSIBLE World web text into transparent daily risk factors and stock-specific betas for quantitative research." url: "https://nosible.com/blog/the-future-that-could-have-been-turning-web-text-into-semantic-stock-betas" --- [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) # The Future That Could Have Been: Turning Web Text into Semantic Stock Betas [Stuart Reid](https://www.linkedin.com/in/stuartgordonreid/) 2026-08-05 11 min read Copy as Markdown ![Thirty semantic factors rendered as daily time series on a dark grid, from geopolitical risk to firm-level political risk.](https://nosible.com/images/2026/08/semantic-factors-hero.png) Markets move on wars, tariffs, pandemics, elections, and central bank decisions. These events are reported in text. The problem? Risk models cannot read. They are built from prices and financial statements because that is what we could measure when they were invented. Today we are releasing the fix. Today we're proud to announce the release of `semantic-factors`, an open-source Python library that turns NOSIBLE World coverage into daily risk factors and stock-level betas. Simply define a risk in plain text. The package quantifies it every day going back more than a decade. All you need is a NOSIBLE World API key. [Explore the 30 semantic factors](https://nosible.com/semantic-factors) or [read the package on GitHub](https://github.com/NosibleAI/semantic-factors). ## Risk models cannot read In 1973, [Robert Merton's ICAPM](https://breesefine7110.tulane.edu/wp-content/uploads/sites/16/2015/10/Merton-Int.-CAPM.pdf) showed that equilibrium expected returns compensate investors for exposure to shocks that change future investment opportunities. What the theory left unresolved was empirical identification: which state variables summarize the information investors actually possess? Contemporary empirical finance largely worked with structured return data, under genuine computational constraints. Fifty years later, news text provides a new measurement layer. As [Bybee, Kelly, and Su](https://doi.org/10.1093/rfs/hhad042) put it, ICAPM risk is tied to news about state variables tracking wealth and future investment opportunities. [Fama and MacBeth's 1973 empirical work](https://www.efalken.com/LowVolClassics/famaMacBeth73.pdf) used structured monthly percentage returns and described the balance between computation costs and the desire to reform portfolios frequently. That was a sensible research design for its time. It also left a clear opening for a new measurement layer built from dated observations of what is happening in the world. Researchers have published more than 400 factors. The [A Census of the Factor Zoo](https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3341728) shows why a factor needs careful definition, transparent construction, and honest validation. Research on [factor decay](https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2156623) shows why signals need to be monitored after publication. The [replication literature](https://academic.oup.com/rfs/article/33/5/2019/5236964) is a reminder that a result is more useful when other researchers can reproduce the construction. Text contains information that prices and filings do not. A news report can describe a military escalation, an export restriction, a new epidemic, or a change in central-bank language before the event has a clean financial representation. The problem is turning that language into a dated, comparable measurement without hiding the research choices or obscuring the researcher's judgment. NOSIBLE World provides the point-in-time event archive. `semantic-factors` provides the definition layer. Together they let a quant researcher ask a direct question: what did the world say about this risk on each day in the archive? ## A worked example: geopolitical risk The gallery below shows the public package-rendered outputs for 30 example semantic factors. The geopolitical-risk definition uses 24 relevance anchors and 3 polarity pairs. The full definition is visible in the [semantic-factors explorer](https://nosible.com/semantic-factors#explorer), where the anchors, polarity sentences, code, research references, CSV, and chart can be inspected together. The following is a complete runnable example using the full geopolitical-risk definition. It uses the same public World search path as the published example and requires a NOSIBLE API key plus an OpenRouter API key to recompute the series. ``` import os from pathlib import Path from semantic_factors import SemanticFactor anchors = [ "Military forces attack, invade, or occupy another sovereign state's territory.", "Fighting expands across new fronts and into populated civilian areas.", "A ceasefire collapses and sustained combat operations resume.", "Air strikes or naval attacks hit another country's territory.", "A government sends troops and weapons across an international border to begin hostilities.", "Armed forces exchange direct fire across a disputed frontier zone.", "An organised group bombs or shoots civilians to cause terror.", "A state-backed group sabotages a cross-border pipeline used by a rival government.", "An armed proxy attacks shipping, airports, or energy facilities abroad.", "Armed attackers seize a civilian district and hold it against national security forces.", "An armed cross-border movement attacks a border post and demands political concessions.", "Security services report a credible plot against major government buildings.", "Two governments exchange threats and move toward an open confrontation.", "A territorial dispute worsens and both claimants reinforce their garrisons.", "Diplomatic talks break down and each government withdraws its ambassador.", "Rival warships or aircraft conduct dangerous manoeuvres near each other.", "A government issues an ultimatum demanding the withdrawal of foreign troops from disputed territory.", "Rival states expel diplomats and suspend their longstanding treaties.", "An attempted coup challenges control of an elected national government.", "Security forces lose control of major cities during mass protests.", "A disputed election leads to violence between rival armed movements.", "Armed factions fight each other for control of national territory.", "A government collapses and no authority can enforce public order.", "Insurgents seize provinces and the national army withdraws from them.", ] poles = [ ( [ "Direct interstate combat is occurring between organized armed forces.", ], [ "No direct interstate combat is occurring between organized armed forces.", ], ), ( [ "An organized political-violence attack against civilians is active.", ], [ "No organized political-violence attack against civilians is active.", ], ), ( [ "Two governments have severed diplomatic relations during an active security dispute.", ], [ "Two governments maintain diplomatic relations during an active security dispute.", ], ), ] factor = SemanticFactor( name="gpr_global", anchors=anchors, poles=poles, filters=None, floor=0.30, aggregate="max", ) output = Path("output") output.mkdir(parents=True, exist_ok=True) factor.compute( weights="netlocs", residualize=False, stabilize=True, polarity_weighted=True, start_from="2015-01-01", stop_at="2026-06-26", api_key=os.environ["NOSIBLE_API_KEY"], embed_api_key=os.environ["OPENROUTER_API_KEY"], ) factor.to_csv( path=str(output / "gpr_global.csv"), ) factor.plot( title="Geopolitical risk", annotate=False, path=str(output / "gpr_global.png"), ) ``` The full geopolitical-risk definition has 24 relevance anchors and 3 polarity pairs. Inspect and edit every sentence in the [semantic-factors explorer](https://nosible.com/semantic-factors#explorer), then replace them with your own research question. ## From a factor to a stock beta Once the factor has been computed, estimate a stock-specific loading by aligning the factor and market data on the same weekly clock: ``` r_i,t = alpha_i + beta_market,i * r_market,t + beta_semantic,i * Delta F_t + epsilon_i,t ``` Here, r_i,t is the stock return, r_market,t is the SPY return, and Delta F_t is the change in the weekly semantic factor. The beta input uses the same 0.25-0.50, k=6 squash as the public charts, followed by the 30-day log1p geometric mean. The seven-day average shown for presentation is not used in the regression because it adds avoidable lag. The calculation uses EODHD adjusted closes, SPY.US as the market control, and standardizes the weekly factor change over the overlapping regression window. The published GPR example has 586 overlapping weekly observations. These are real historical OLS outputs from the public example, not placeholder values: | Stock | Role | GPR semantic beta | | --- | --- | --- | | RTX.US | defense | +0.110% | | LMT.US | defense | +0.199% | | CAT.US | industrial | +0.027% | | DAL.US | airline | +0.066% | These values use one signal definition across every stock. The defense names have positive loadings in this sample, while the industrial and airline estimates are smaller. That is the kind of cross-sectional difference the factor is intended to expose, not a hand-written interpretation of one company. A positive loading means the stock tended to have a higher return when the standardized weekly factor change was positive, after controlling for the market in this specification. A negative loading means the opposite association. A beta is a conditional historical association, not a claim that the factor caused the return or that the relationship will persist. The [semantic-factors page](https://nosible.com/semantic-factors) lets you switch between all 30 factors and inspect their company beta outputs. The numbers belong to a measurement window, model specification, and data source. They are inputs for research, not investment recommendations or promises about future performance. This is not a one-to-one reproduction of the published GPR methodology. The corpus, event definitions, anchors, threshold, weighting, polarity construction, and aggregation are different. The point is to make a new semantic factor possible, inspectable, and reproducible from the World data. ## See the 30 semantic factors These are package-rendered outputs from the current public snapshot, not mock data. Click any chart to open the full-resolution image. The same definitions, code, and downloadable data are available in the [semantic-factors explorer](https://nosible.com/semantic-factors#explorer). [![Geopolitical risk semantic factor chart](https://nosible.com/data/semantic-factors/plots/gpr_global.png?v=20260805) Geopolitical risk semantic factor chart 01 / Geopolitical risk](https://nosible.com/data/semantic-factors/plots/gpr_global.png?v=20260805) [![Trade policy uncertainty semantic factor chart](https://nosible.com/data/semantic-factors/plots/tpu_us.png?v=20260805) Trade policy uncertainty semantic factor chart 02 / Trade policy uncertainty](https://nosible.com/data/semantic-factors/plots/tpu_us.png?v=20260805) [![Pandemic and health risk semantic factor chart](https://nosible.com/data/semantic-factors/plots/pandemic_health_us.png?v=20260805) Pandemic and health risk semantic factor chart 03 / Pandemic and health risk](https://nosible.com/data/semantic-factors/plots/pandemic_health_us.png?v=20260805) [![Sanctions and export controls semantic factor chart](https://nosible.com/data/semantic-factors/plots/sanctions_us.png?v=20260805) Sanctions and export controls semantic factor chart 04 / Sanctions and export controls](https://nosible.com/data/semantic-factors/plots/sanctions_us.png?v=20260805) [![Recession nowcast semantic factor chart](https://nosible.com/data/semantic-factors/plots/recession_nowcast_us.png?v=20260805) Recession nowcast semantic factor chart 05 / Recession nowcast](https://nosible.com/data/semantic-factors/plots/recession_nowcast_us.png?v=20260805) [![Oil supply risk semantic factor chart](https://nosible.com/data/semantic-factors/plots/oil_supply_risk.png?v=20260805) Oil supply risk semantic factor chart 06 / Oil-supply risk](https://nosible.com/data/semantic-factors/plots/oil_supply_risk.png?v=20260805) [![Supply chain pressure semantic factor chart](https://nosible.com/data/semantic-factors/plots/supply_chain_pressure.png?v=20260805) Supply chain pressure semantic factor chart 07 / Supply-chain pressure](https://nosible.com/data/semantic-factors/plots/supply_chain_pressure.png?v=20260805) [![Economic policy uncertainty semantic factor chart](https://nosible.com/data/semantic-factors/plots/epu_us.png?v=20260805) Economic policy uncertainty semantic factor chart 08 / Economic policy uncertainty](https://nosible.com/data/semantic-factors/plots/epu_us.png?v=20260805) [![Financial stress semantic factor chart](https://nosible.com/data/semantic-factors/plots/financial_stress_global.png?v=20260805) Financial stress semantic factor chart 09 / Financial stress](https://nosible.com/data/semantic-factors/plots/financial_stress_global.png?v=20260805) [![Energy security semantic factor chart](https://nosible.com/data/semantic-factors/plots/energy_security_us.png?v=20260805) Energy security semantic factor chart 10 / Energy security](https://nosible.com/data/semantic-factors/plots/energy_security_us.png?v=20260805) [![Inflation attention semantic factor chart](https://nosible.com/data/semantic-factors/plots/inflation_attention_us_current.png?v=20260805) Inflation attention semantic factor chart 11 / Inflation attention](https://nosible.com/data/semantic-factors/plots/inflation_attention_us_current.png?v=20260805) [![Climate policy uncertainty semantic factor chart](https://nosible.com/data/semantic-factors/plots/climate_policy_uncertainty_us.png?v=20260805) Climate policy uncertainty semantic factor chart 12 / Climate-policy uncertainty](https://nosible.com/data/semantic-factors/plots/climate_policy_uncertainty_us.png?v=20260805) [![Bank regulation uncertainty semantic factor chart](https://nosible.com/data/semantic-factors/plots/bank_reg_uncertainty_us.png?v=20260805) Bank regulation uncertainty semantic factor chart 13 / Bank-regulation uncertainty](https://nosible.com/data/semantic-factors/plots/bank_reg_uncertainty_us.png?v=20260805) [![Fiscal uncertainty semantic factor chart](https://nosible.com/data/semantic-factors/plots/fiscal_uncertainty_us.png?v=20260805) Fiscal uncertainty semantic factor chart 14 / Fiscal uncertainty](https://nosible.com/data/semantic-factors/plots/fiscal_uncertainty_us.png?v=20260805) [![Country geopolitical risk semantic factor chart](https://nosible.com/data/semantic-factors/plots/gpr_us.png?v=20260805) Country geopolitical risk semantic factor chart 15 / Country geopolitical risk](https://nosible.com/data/semantic-factors/plots/gpr_us.png?v=20260805) [![Bilateral tension semantic factor chart](https://nosible.com/data/semantic-factors/plots/bilateral_tension_us.png?v=20260805) Bilateral tension semantic factor chart 16 / Bilateral tension](https://nosible.com/data/semantic-factors/plots/bilateral_tension_us.png?v=20260805) [![Nuclear threat semantic factor chart](https://nosible.com/data/semantic-factors/plots/nuclear_threat_us.png?v=20260805) Nuclear threat semantic factor chart 17 / Nuclear threat](https://nosible.com/data/semantic-factors/plots/nuclear_threat_us.png?v=20260805) [![Food security semantic factor chart](https://nosible.com/data/semantic-factors/plots/food_security_us.png?v=20260805) Food security semantic factor chart 18 / Food security](https://nosible.com/data/semantic-factors/plots/food_security_us.png?v=20260805) [![Commodity supply risk semantic factor chart](https://nosible.com/data/semantic-factors/plots/commodity_supply_risk_us.png?v=20260805) Commodity supply risk semantic factor chart 19 / Commodity supply risk](https://nosible.com/data/semantic-factors/plots/commodity_supply_risk_us.png?v=20260805) [![Sovereign distress semantic factor chart](https://nosible.com/data/semantic-factors/plots/sovereign_distress_us.png?v=20260805) Sovereign distress semantic factor chart 20 / Sovereign distress](https://nosible.com/data/semantic-factors/plots/sovereign_distress_us.png?v=20260805) [![Risk on risk off semantic factor chart](https://nosible.com/data/semantic-factors/plots/risk_off_global.png?v=20260805) Risk on risk off semantic factor chart 21 / Risk-on / risk-off](https://nosible.com/data/semantic-factors/plots/risk_off_global.png?v=20260805) [![Migration policy semantic factor chart](https://nosible.com/data/semantic-factors/plots/migration_policy_us.png?v=20260805) Migration policy semantic factor chart 22 / Migration policy](https://nosible.com/data/semantic-factors/plots/migration_policy_us.png?v=20260805) [![Currency crisis semantic factor chart](https://nosible.com/data/semantic-factors/plots/currency_crisis_usd.png?v=20260805) Currency crisis semantic factor chart 23 / Currency crisis](https://nosible.com/data/semantic-factors/plots/currency_crisis_usd.png?v=20260805) [![Cyber risk semantic factor chart](https://nosible.com/data/semantic-factors/plots/cyber_risk_us.png?v=20260805) Cyber risk semantic factor chart 24 / Cyber risk](https://nosible.com/data/semantic-factors/plots/cyber_risk_us.png?v=20260805) [![Terrorism and unrest semantic factor chart](https://nosible.com/data/semantic-factors/plots/terror_unrest_us.png?v=20260805) Terrorism and unrest semantic factor chart 25 / Terrorism and unrest](https://nosible.com/data/semantic-factors/plots/terror_unrest_us.png?v=20260805) [![Partisan conflict semantic factor chart](https://nosible.com/data/semantic-factors/plots/partisan_conflict_us.png?v=20260805) Partisan conflict semantic factor chart 26 / Partisan conflict](https://nosible.com/data/semantic-factors/plots/partisan_conflict_us.png?v=20260805) [![Monetary policy uncertainty semantic factor chart](https://nosible.com/data/semantic-factors/plots/mpu_fed.png?v=20260805) Monetary policy uncertainty semantic factor chart 27 / Monetary policy uncertainty](https://nosible.com/data/semantic-factors/plots/mpu_fed.png?v=20260805) [![Equity volatility tracker semantic factor chart](https://nosible.com/data/semantic-factors/plots/equity_volatility_us.png?v=20260805) Equity volatility tracker semantic factor chart 28 / Equity volatility tracker](https://nosible.com/data/semantic-factors/plots/equity_volatility_us.png?v=20260805) [![Firm level political risk semantic factor chart](https://nosible.com/data/semantic-factors/plots/firm_political_risk_us.png?v=20260805) Firm level political risk semantic factor chart 29 / Firm-level political risk](https://nosible.com/data/semantic-factors/plots/firm_political_risk_us.png?v=20260805) [![Hawk dove sentiment semantic factor chart](https://nosible.com/data/semantic-factors/plots/fed_hawk_dove.png?v=20260805) Hawk dove sentiment semantic factor chart 30 / Hawk-dove sentiment](https://nosible.com/data/semantic-factors/plots/fed_hawk_dove.png?v=20260805) Download the [combined 30-factor daily CSV](https://nosible.com/data/semantic-factors/nosible-macro-risk-indicators-daily.csv), inspect the [full semantic-factor page](https://nosible.com/semantic-factors), or open the [package repository on GitHub](https://github.com/NosibleAI/semantic-factors). ## What you can do with it With `semantic-factors`, a quant team can: - monitor a risk vendors do not cover in the language your team uses; - measure a developing exposure before it appears in a conventional risk report; - build differentiated signals from a transparent research definition; - compare the same concept across countries, sectors, or event scopes; - rerun the measurement when the question changes without waiting for vendors. The library is useful for research that begins with a question rather than a ticker. What does escalation look like in the news? How does a supply shock appear in the text? Which companies have historically moved with the resulting time series? ## Why not just ask an LLM? An answer from an LLM is not a dated time series. It is difficult to reproduce, difficult to audit, and usually disconnected from a fixed historical corpus. `semantic-factors` uses a point-in-time World archive, explicit sentences, a declared threshold, and a deterministic aggregation path. The output can be inspected event by event. The definition can be versioned. The CSV can be downloaded. The chart can be regenerated. A colleague can challenge an anchor or a polarity pair and rerun the measurement. This makes the system useful for research workflows where the question matters as much as the number. It does not remove judgment. It puts judgment where a researcher can see, test, challenge, and carefully revise it. ## Get started Install the open-source package: ``` python -m pip install "semantic-factors[plot]" ``` To recompute a factor, you need access to NOSIBLE World and an embedding provider: ``` export NOSIBLE_API_KEY=nos_sk_... export OPENROUTER_API_KEY=sk-or-... ``` The public page includes a snapshot of the 30 worked examples so you can inspect the definitions and outputs without recomputing them. When you are ready to build your own, start with a concept, write the sentences, and run the package. [Explore the 30 semantic factors](https://nosible.com/semantic-factors) [Start a NOSIBLE trial](https://nosible.com/start-trial) [View the package on GitHub](https://github.com/NosibleAI/semantic-factors) For fifty years, quantitative researchers have turned observations into factors. Now the observation can begin with the words describing what is happening in the world. [All Research](https://nosible.com/blog) Related Research ![Signed contrast number line separating systemic and idiosyncratic risk](https://nosible.com/images/2026/06/walkthrough_step9_number_line.png) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ### [Two Tricks for Turning Sentence Embeddings into Clean Features](https://nosible.com/blog/the-contrastive-geometry-of-risk) 2026-06-18 14 min read ![Daily NOSIBLE Trade Policy Uncertainty index compared with the published index](https://nosible.com/images/2026/06/nosible-tpu-daily.png) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ### [An Embedding-Based Approach to Trade and Economic Policy Uncertainty](https://nosible.com/blog/an-embedding-based-approach-to-trade-and-economic-policy-uncertainty) 2026-06-17 23 min read ![High-resolution chart of selected 35B-A3B serving records across a 350-plus experiment campaign, from Qwen 3.5 FP8 on H200 at $0.218400 per million completion tokens in Experiment 006 to Qwen 3.6 source NVFP4 at $0.125959 in Experiment 332](https://nosible.com/images/2026/08/qwen-35b-a3b-serving-campaign-hero.svg) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) ### [What 350+ Experiments Taught Us About Serving Qwen 3.6 35B-A3B](https://nosible.com/blog/what-350-experiments-taught-us-about-serving-qwen-3-6) 2026-08-17 15 min read > Learn how semantic-factors turns NOSIBLE World web text into transparent daily risk factors and stock-specific betas for quantitative research. **URL:** https://nosible.com/blog/the-future-that-could-have-been-turning-web-text-into-semantic-stock-betas --- --- title: "Web Search" description: "Research on building, operating, and improving web-scale search systems for AI and quantitative workflows." url: "https://nosible.com/blog/tag/web-search" --- /res Research # Web Search Research on building, operating, and improving web-scale search systems for AI and quantitative workflows. [All](https://nosible.com/blog) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Financial Sentiment](https://nosible.com/blog/tag/financial-sentiment) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [Trading Signals](https://nosible.com/blog/tag/trading-signals) [Web Search](https://nosible.com/blog/tag/web-search) 5 articles ![NOSIBLE World knowledge graph showing entity connections over a decade](https://nosible.com/images/2026/07/kg-hero-decade.png) [Web Search](https://nosible.com/blog/tag/web-search) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ## [Point-in-Time Knowledge Graphs over Named Entities with NOSIBLE World](https://nosible.com/blog/point-in-time-knowledge-graphs-over-named-entities) Companies change: people leave, products die, mergers happen. This post shows how to build point-in-time knowledge graphs over any company from news alone, using NOSIBLE World, which covers 38 thousand tickers, 3.2 million organizations and 6.4 million people. Named entity recognition supplies the nodes, lift-scored co-mention supplies the links, and the document date keeps every yearly view safe to backtest. 2026-07-16 8 min read ![Abstract vortex illustration representing self-organizing web-scale search facets](https://nosible.com/blog/illustrations/vortex.png) [Web Search](https://nosible.com/blog/tag/web-search) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) ## [Can Faceted Search at Web-Scale Self Organize?](https://nosible.com/blog/can-faceted-search-at-web-scale-self-organize) Can Faceted Search at Web-Scale Self Organize? As it turns out, yes it can! In this post we outline our new and improved adaptive named entity tagging system! 2025-10-16 7 min read ![Cybernaut-1 illustration representing agentic search with Monte Carlo Tree Search](https://nosible.com/blog/illustrations/cyber.png) [Web Search](https://nosible.com/blog/tag/web-search) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) ## [Introducing Cybernaut-1: Agentic Search using MCTS](https://nosible.com/blog/introducing-cybernaut-1-agentic-search-with-mcts) Cybernaut-1 combines our powerful hybrid-3 search algorithm with LLM-guided Monte Carlo Tree Search to deliver world class search results on difficult queries. 2025-08-26 2 min read ![Railway-track illustration representing the road to Cybernaut-1](https://nosible.com/blog/illustrations/track.png) [Web Search](https://nosible.com/blog/tag/web-search) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) ## [The Road to Cybernaut-1: Rebuilding Search for AI](https://nosible.com/blog/the-road-to-cybernaut-1) AI needs its own search engine. This is how we’re rebuilding search for AI -- and the road to Cybernaut-1, the first high-trust agentic search engine. 2025-08-20 17 min read ![Magnifying-glass illustration representing vector search across company news](https://nosible.com/blog/illustrations/inspect.png) [Web Search](https://nosible.com/blog/tag/web-search) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) [Trading Signals](https://nosible.com/blog/tag/trading-signals) ## [Using Vector Search to See Signals in Company News](https://nosible.com/blog/using-vector-search-to-see-signals-in-company-news) How we use vector search to extract investment signals from a multi-terabyte company news dataset that currently contains over 55 million embeddings, 150+ million sentences, 4+ billion words, and 5+ billion GPT tokens. 2024-01-21 21 min read [All Research](https://nosible.com/blog) > Research on building, operating, and improving web-scale search systems for AI and quantitative workflows. **URL:** https://nosible.com/blog/tag/web-search --- --- title: "Information Retrieval" description: "Research on finding, routing, ranking, and structuring evidence from the web for reliable retrieval." url: "https://nosible.com/blog/tag/information-retrieval" --- /res Research # Information Retrieval Research on finding, routing, ranking, and structuring evidence from the web for reliable retrieval. [All](https://nosible.com/blog) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Financial Sentiment](https://nosible.com/blog/tag/financial-sentiment) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [Trading Signals](https://nosible.com/blog/tag/trading-signals) [Web Search](https://nosible.com/blog/tag/web-search) 5 articles ![NOSIBLE World knowledge graph showing entity connections over a decade](https://nosible.com/images/2026/07/kg-hero-decade.png) [Web Search](https://nosible.com/blog/tag/web-search) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ## [Point-in-Time Knowledge Graphs over Named Entities with NOSIBLE World](https://nosible.com/blog/point-in-time-knowledge-graphs-over-named-entities) Companies change: people leave, products die, mergers happen. This post shows how to build point-in-time knowledge graphs over any company from news alone, using NOSIBLE World, which covers 38 thousand tickers, 3.2 million organizations and 6.4 million people. Named entity recognition supplies the nodes, lift-scored co-mention supplies the links, and the document date keeps every yearly view safe to backtest. 2026-07-16 8 min read ![Abstract vortex illustration representing self-organizing web-scale search facets](https://nosible.com/blog/illustrations/vortex.png) [Web Search](https://nosible.com/blog/tag/web-search) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) ## [Can Faceted Search at Web-Scale Self Organize?](https://nosible.com/blog/can-faceted-search-at-web-scale-self-organize) Can Faceted Search at Web-Scale Self Organize? As it turns out, yes it can! In this post we outline our new and improved adaptive named entity tagging system! 2025-10-16 7 min read ![Cybernaut-1 illustration representing agentic search with Monte Carlo Tree Search](https://nosible.com/blog/illustrations/cyber.png) [Web Search](https://nosible.com/blog/tag/web-search) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) ## [Introducing Cybernaut-1: Agentic Search using MCTS](https://nosible.com/blog/introducing-cybernaut-1-agentic-search-with-mcts) Cybernaut-1 combines our powerful hybrid-3 search algorithm with LLM-guided Monte Carlo Tree Search to deliver world class search results on difficult queries. 2025-08-26 2 min read ![Railway-track illustration representing the road to Cybernaut-1](https://nosible.com/blog/illustrations/track.png) [Web Search](https://nosible.com/blog/tag/web-search) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) ## [The Road to Cybernaut-1: Rebuilding Search for AI](https://nosible.com/blog/the-road-to-cybernaut-1) AI needs its own search engine. This is how we’re rebuilding search for AI -- and the road to Cybernaut-1, the first high-trust agentic search engine. 2025-08-20 17 min read ![Magnifying-glass illustration representing vector search across company news](https://nosible.com/blog/illustrations/inspect.png) [Web Search](https://nosible.com/blog/tag/web-search) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) [Trading Signals](https://nosible.com/blog/tag/trading-signals) ## [Using Vector Search to See Signals in Company News](https://nosible.com/blog/using-vector-search-to-see-signals-in-company-news) How we use vector search to extract investment signals from a multi-terabyte company news dataset that currently contains over 55 million embeddings, 150+ million sentences, 4+ billion words, and 5+ billion GPT tokens. 2024-01-21 21 min read [All Research](https://nosible.com/blog) > Research on finding, routing, ranking, and structuring evidence from the web for reliable retrieval. **URL:** https://nosible.com/blog/tag/information-retrieval --- --- title: "Point-in-Time Knowledge Graphs over Named Entities with NOSIBLE World" description: "Build point-in-time knowledge graphs from dated NOSIBLE World events using named entities and lift-scored co-mentions—without lookahead bias." url: "https://nosible.com/blog/point-in-time-knowledge-graphs-over-named-entities" --- [Web Search](https://nosible.com/blog/tag/web-search) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) # Point-in-Time Knowledge Graphs over Named Entities with NOSIBLE World [Stuart Reid](https://www.linkedin.com/in/stuartgordonreid/) 2026-07-16 8 min read Copy as Markdown ![NOSIBLE World knowledge graph showing entity connections over a decade](https://nosible.com/images/2026/07/kg-hero-decade.png) Companies change over time. People leave. Products evolve. Mergers happen. One of the topics we are interested in at NOSIBLE is capturing and structuring point-in-time knowledge. In this post we demonstrate how you can build point-in-time knowledge graphs from NOSIBLE World. For reference, World covers 38 thousand tickers, 3.2 million organizations, and 6.4 million people. ## The Ideas in Brief Named entity recognition is a computer program that automatically reads text to find and categorize important nouns like people, places, and companies. A sentence like "NVIDIA unveiled Blackwell at GTC in San Jose, and Jensen Huang demonstrated it against AMD" carries four kinds of entity: an organization (NVIDIA, AMD), a product (Blackwell), a person (Jensen Huang), and a place (San Jose). NOSIBLE World runs NER over every event and records many entity types; this post focuses on five: ORG, PERSON, GPE and LOC for places, and PRODUCT. A knowledge graph is a web linking important concepts like people and companies based on how often they are mentioned together. The nodes in the knowledge graph are the entities our NER found. The edges between them are inferred from co-mentions. For example, when NVIDIA (ORG) is co-mentioned often enough with Blackwell (PRODUCT), those two nodes are connected using an edge. The strength of the connection between two nodes is proportional to the frequency-adjusted probability of them co-occurring. Point-in-time means that every connection in a knowledge graph was knowable at that specific moment. A 2018 view, for example, is built entirely from information available during that year to ensure the graph remains safe for backtesting. When future knowledge seeps into the past, it creates lookahead bias. For instance, if a model evaluating 2007 data already knows the newly announced iPhone will become a massive commercial success, its predictions rely on impossible foresight rather than historical reality. ## How It Works We treat edge formation as a link discovery problem. For a company, a year and a candidate entity, we count how many events co-mention them, then compare that against how often it appears across all news that year. That comparison is lift: ``` lift = (co / target_events) / (entity_events / total_events) ``` Lift of 1.0 is chance. Higher means the pair co-occurs more than chance would predict. Ubiquitous entities like "United States" appear everywhere, so their lift against any one company is near 1.0 and they fall away. A specific pairing like NVIDIA and Blackwell scores high: Blackwell was co-mentioned with NVIDIA in 69 of NVIDIA's 2,651 events in 2024 and appeared in only 75 events in the whole corpus that year, which works out to roughly 780 times chance. Raw co-mention counts alone would just surface whatever is loudest, so lift does the real work, and a normalized version of it lets a strong link on a rare product rank alongside one on a common name. Before ranking we tidy the entities: surface forms are canonicalized, so Apple, Apple Inc. and apple collapse to one node and a lone surname folds into the fuller name kept that year. The point-in-time property is enforced by bucketing every edge on its document date and scoring it against the same year's background, with no cumulative or forward-looking state. Sampling edges at random and pulling the underlying source documents back from the database, all of them fall inside the year the edge belongs to: 24 of 24 in the last run, across 955 real documents. NOSIBLE World supplies the raw material. It turns the article stream into one record per real event, however many outlets cover it, and each record carries a date, its NER entities, resolved company tickers, a coverage count of distinct publishers, and an embedding of its meaning. For this proof of concept we scanned every daily slice from 2015 to mid 2026 and pulled the events that mention NVIDIA, Apple or Coca-Cola: 88,359 in all. Across the period that resolves to 332 distinct products for NVIDIA and 274 for Apple, along with hundreds of people, organizations and places for each company. The matrices below show only the strongest of them. Here is one such record, the January 2022 Microsoft and Activision Blizzard deal, abridged to the fields this post uses. The full event also carries an embedding of its meaning, a twenty-facet ontology (GICS, IPTC and NOSIBLE event types among them), per-publisher provenance, and identifiers for each ticker (ISIN, LEI, FIGI), all left out here for length. ``` { "event": { "date": "2022-01-18", "title": "Microsoft Acquires Activision Blizzard in $68.7 Billion Gaming Deal", "country": "United States" }, "signals": { "sentiment": "positive", "materiality_score": 0.83 }, "coverage": { "total_coverage": 2740, "total_netlocs": 2392 }, "entities": { "PERSON": { "Bobby Kotick": 22, "Phil Spencer": 15, "Satya Nadella": 6 }, "ORG": { "Microsoft": 202, "Activision Blizzard": 157, "Xbox": 26 }, "GPE": { "United States": 3, "California": 3 }, "PRODUCT": { "Call of Duty": 20, "Xbox Game Pass": 6, "Playstation": 3 } }, "tickers": [ { "name": "Microsoft Corporation", "ticker_eodhd": "MSFT.US" }, { "name": "Activision Blizzard Inc", "ticker_eodhd": "ATVI.US" }, { "name": "Sony Corp", "ticker_eodhd": "6758.TSE" } ] } ``` ## Worked Examples We show each company the same way: first its force-directed network for one year, where nodes are the entities it was co-mentioned with and lines join entities that also co-occur with each other, so related nodes cluster; then the matrix that best tells its story over time. The networks colour nodes by category and the matrices colour by trend, green where a share grew, white where it held, red where it shrank. ### NVIDIA NVIDIA's 2024 network splits into the groups the news runs together: the memory and foundry supply chain, the accelerator platform of GeForce, GPU and Blackwell, and the AI cloud of Microsoft, Google and OpenAI, with AMD and Intel as the peers. ![NVIDIA's 2024 co-mention network. Nodes are entities co-mentioned with NVIDIA; lines join entities that co-occur with each other, so related nodes cluster. Brand green marks products, lighter green peers, mist people, muted green organizations, an outline places; size is association strength. The supply chain, the accelerator platform and the AI cloud each form a group. Built from news on or before 31 December 2024.](https://nosible.com/images/2026/07/kg-network-nvda.png) NVIDIA's 2024 co-mention network. Nodes are entities co-mentioned with NVIDIA; lines join entities that co-occur with each other, so related nodes cluster. Brand green marks products, lighter green peers, mist people, muted green organizations, an outline places; size is association strength. The supply chain, the accelerator platform and the AI cloud each form a group. Built from news on or before 31 December 2024. The product matrix then shows the roadmap arriving on cue, each product entering the year the news starts to discuss it: RTX in 2018, DLSS in 2019, the A100 in 2020, then Blackwell, HBM and the H200 in 2024. ![The NVIDIA product line-up over time, showing 40 of the 332 products NVIDIA is co-mentioned with. Dot size is that year's association strength; colour is how the product's share of NVIDIA's coverage changed from the year before, green for growing, red for shrinking, neutral for flat. RTX enters in 2018, the A100 in 2020, and Blackwell, HBM and the H200 arrive in 2024.](https://nosible.com/images/2026/07/kg-evolution-matrix.png) The NVIDIA product line-up over time, showing 40 of the 332 products NVIDIA is co-mentioned with. Dot size is that year's association strength; colour is how the product's share of NVIDIA's coverage changed from the year before, green for growing, red for shrinking, neutral for flat. RTX enters in 2018, the A100 in 2020, and Blackwell, HBM and the H200 arrive in 2024. ### Apple Apple's 2024 network centres on its own product line, the iPhone, iPad, the Pro and Pro Max and Vision Pro, ringed by peers, the people who cover it, and the platforms it competes and partners with. ![Apple's 2024 co-mention network, drawn the same way. Its product line sits at the centre, ringed by peers such as Samsung and Foxconn, the analysts and executives who cover it (Tim Cook, Mark Gurman, Ming-Chi Kuo), and the platforms it competes and partners with (Google, Microsoft, Spotify). Built from news on or before 31 December 2024.](https://nosible.com/images/2026/07/kg-network-aapl.png) Apple's 2024 co-mention network, drawn the same way. Its product line sits at the centre, ringed by peers such as Samsung and Foxconn, the analysts and executives who cover it (Tim Cook, Mark Gurman, Ming-Chi Kuo), and the platforms it competes and partners with (Google, Microsoft, Spotify). Built from news on or before 31 December 2024. The product matrix traces the arc over time: the iPhone and iPad run throughout, the Apple Watch and AirPods arrive in the mid 2010s, the Pro Max line follows from 2019, then Vision Pro in 2025. ![Apple's product line-up over time, 40 of 274 products, drawn the same way. The iPhone, iPad and App Store run throughout; the Apple Watch and AirPods arrive in the mid 2010s, the iPhone X in 2018, AirPods Pro and the Pro Max line from 2019, then AirTag and Dynamic Island, and Vision Pro in 2025.](https://nosible.com/images/2026/07/kg-aapl-products-matrix.png) Apple's product line-up over time, 40 of 274 products, drawn the same way. The iPhone, iPad and App Store run throughout; the Apple Watch and AirPods arrive in the mid 2010s, the iPhone X in 2018, AirPods Pro and the Pro Max line from 2019, then AirTag and Dynamic Island, and Vision Pro in 2025. ### Coca-Cola Coca-Cola's 2021 network is its own world: the bottling system on one side, the beverage competitors from PepsiCo to Anheuser-Busch on another. The lower cluster is a single news story, the backlash to Georgia's 2021 voting law, when Coca-Cola and Delta criticized it and Major League Baseball moved its All-Star Game out of Atlanta, which is why Georgia, Atlanta, Delta and the MLB sit together. ![Coca-Cola's 2021 co-mention network, drawn the same way. The bottling system clusters on one side and the beverage competitors on another. The lower group is the backlash to Georgia's 2021 voting law, when Coca-Cola and Delta criticized it and Major League Baseball pulled its All-Star Game from Atlanta, so Georgia, Atlanta, Delta and the MLB appear together. Built from news on or before 31 December 2021.](https://nosible.com/images/2026/07/kg-network-ko.png) Coca-Cola's 2021 co-mention network, drawn the same way. The bottling system clusters on one side and the beverage competitors on another. The lower group is the backlash to Georgia's 2021 voting law, when Coca-Cola and Delta criticized it and Major League Baseball pulled its All-Star Game from Atlanta, so Georgia, Atlanta, Delta and the MLB appear together. Built from news on or before 31 December 2021. Its best story is in people. The chief-executive handover from Muhtar Kent to James Quincey reads straight off the news, and Mark Zuckerberg appears once, in 2020, from the advertising boycott Coca-Cola led. ![Coca-Cola's associated people over time. Muhtar Kent is present from 2015 to 2018, James Quincey from 2016 on, and Cristiano Ronaldo shows up as a single spike in 2021, the year of his Euro 2020 press-conference moment.](https://nosible.com/images/2026/07/kg-ko-people-matrix.png) Coca-Cola's associated people over time. Muhtar Kent is present from 2015 to 2018, James Quincey from 2016 on, and Cristiano Ronaldo shows up as a single spike in 2021, the year of his Euro 2020 press-conference moment. ## Run It on Any Company The same pipeline runs over any company with a ticker in NOSIBLE World. Pick a company, choose the entity layers that matter, products for a roadmap, people for the cast, peers for the competitive set, and you get a point-in-time knowledge graph that is safe to backtest, because every edge is stamped with the date it became knowable. NOSIBLE World turns the world's news into a structured, multilingual, de-duplicated event database, one record per real event, with dates, entities, tickers and embeddings already resolved. If you want access to the data, or a graph like this built for your own universe, [start a trial](https://nosible.com/start-trial) or explore the live data at [nosible.world](https://nosible.world/). ## Appendix: Build It Yourself The core of the pipeline is two short functions and a scoring line, shown here as one runnable script. It starts from an event as NOSIBLE World stores it, with entities grouped by type and a mention count per name; counts co-mentions per year, bucketed by the event date so the view is point-in-time by construction; and scores a candidate edge by lift, reproducing the NVIDIA and Blackwell pairing from earlier. ``` import collections def co_mentions_by_year( events: list[dict], target_ticker: str, entity_type: str ) -> dict: """ Count, per year, the events mentioning both a target company and each entity. :param events: NOSIBLE World events, each with a date, tickers and entities. :param target_ticker: EODHD ticker of the company at the centre of the graph. :param entity_type: NER layer to read, for example "PRODUCT" or "PERSON". :return: Mapping of year to a Counter of entity name to co-mention count. """ per_year = collections.defaultdict(collections.Counter) for event in events: tickers = {ticker["ticker_eodhd"] for ticker in event["tickers"]} if target_ticker not in tickers: continue year = event["date"][:4] for name in event["entities"].get(entity_type, {}): per_year[year][name] += 1 return per_year def lift( co: int, target_events: int, entity_events: int, total_events: int ) -> float: """ Association strength of a co-mention versus chance in one period. :param co: Events mentioning both the target and the entity. :param target_events: Events mentioning the target. :param entity_events: Events mentioning the entity, the background. :param total_events: All events in the period. :return: Lift, where 1.0 is chance and higher is a stronger association. """ return (co / target_events) / (entity_events / total_events) # A simplified event. The real NOSIBLE World record (shown earlier) nests the date # under an "event" object and also carries signals, coverage, an embedding and an # ontology; this toy keeps only what the two functions read. event = { "date": "2024-03-18", "tickers": [{"ticker_eodhd": "NVDA.US", "name": "NVIDIA"}], "entities": { "PRODUCT": {"Blackwell": 7, "GPU": 4}, "PERSON": {"Jensen Huang": 3}, "ORG": {"AMD": 2}, "GPE": {"Taiwan": 2} } } # Count co-mentions per year, bucketed by the event date, so the view is # point-in-time by construction. per_year = co_mentions_by_year( events=[event], target_ticker="NVDA.US", entity_type="PRODUCT" ) # Score the NVIDIA and Blackwell edge for 2024 by lift, the pairing from above. # The result is about 780: the pair co-occurs roughly 780 times more than chance # in 2024, so the edge is kept. score = lift( co=69, target_events=2651, entity_events=75, total_events=2249039 ) ``` [All Research](https://nosible.com/blog) Related Research ![Abstract vortex illustration representing self-organizing web-scale search facets](https://nosible.com/blog/illustrations/vortex.png) [Web Search](https://nosible.com/blog/tag/web-search) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) ### [Can Faceted Search at Web-Scale Self Organize?](https://nosible.com/blog/can-faceted-search-at-web-scale-self-organize) 2025-10-16 7 min read ![Cybernaut-1 illustration representing agentic search with Monte Carlo Tree Search](https://nosible.com/blog/illustrations/cyber.png) [Web Search](https://nosible.com/blog/tag/web-search) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) ### [Introducing Cybernaut-1: Agentic Search using MCTS](https://nosible.com/blog/introducing-cybernaut-1-agentic-search-with-mcts) 2025-08-26 2 min read ![Railway-track illustration representing the road to Cybernaut-1](https://nosible.com/blog/illustrations/track.png) [Web Search](https://nosible.com/blog/tag/web-search) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) ### [The Road to Cybernaut-1: Rebuilding Search for AI](https://nosible.com/blog/the-road-to-cybernaut-1) 2025-08-20 17 min read > Build point-in-time knowledge graphs from dated NOSIBLE World events using named entities and lift-scored co-mentions—without lookahead bias. **URL:** https://nosible.com/blog/point-in-time-knowledge-graphs-over-named-entities --- --- title: "Two Tricks for Turning Sentence Embeddings into Clean Features" description: "Turn OpenAI sentence embeddings into clean geographic and systemic-risk features with two deterministic, training-free scoring techniques." url: "https://nosible.com/blog/the-contrastive-geometry-of-risk" --- [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) # Two Tricks for Turning Sentence Embeddings into Clean Features [Stuart Reid](https://www.linkedin.com/in/stuartgordonreid/) 2026-06-18 14 min read Copy as Markdown ![Signed contrast number line separating systemic and idiosyncratic risk](https://nosible.com/images/2026/06/walkthrough_step9_number_line.png) NOSIBLE is a search engine, so naturally we've spent a lot of time staring at vectors. Today we are going to share two tricks that can be used to amplify signal and squash noise. Let's imagine you are working with NOSIBLE World and you want to organize events into three geographic buckets: local, national, or global, then classify whether they are systemic (about the whole market) or idiosyncratic (about specific companies). World does not come with these dimensions, so how would you solve this? You'd reach out to an LLM right? That would work but it's slow, expensive, and massively increases the risk of foreknowledge bias. Today I'm going to show you how you can get this done using nothing more than geometry. We will carry one event through the whole pipeline. > **Input sentence:** *"Central banks across every major economy warn that a synchronized financial shock is freezing credit and driving stock markets lower worldwide."* It should come out **global** and **systemic**. Here is how, in ten steps. ## Step 1: write reference texts for your geographic buckets A bucket is defined by a set of example sentences. We wrote twenty per bucket as near mirror images of one another: the same declarative frames, repeated across all three sets, with only the scale phrase changing from one place, to a whole country, to the whole world. Holding the topic and the sentence style constant means they cancel across the sets, and the only thing left to separate one set from another is geographic scale. Here are the three sets we authored for this post. | Local | National | Global | | --- | --- | --- | | A worsening crisis is now affecting a single town. | A worsening crisis is now affecting the whole country. | A worsening crisis is now affecting the entire world. | | Disruption is being felt across one neighborhood. | Disruption is being felt across the entire nation. | Disruption is being felt across every continent. | | The impact has spread across a single local community. | The impact has spread across the entire country. | The impact has spread across countries around the globe. | | Day by day, the emergency is engulfing one city. | Day by day, the emergency is engulfing the whole nation. | Day by day, the emergency is engulfing nations worldwide. | | Damage now reaches a single district. | Damage now reaches the country from coast to coast. | Damage now reaches the whole planet. | | What began as a small problem is spreading across one village. | What began as a small problem is spreading across the country as a whole. | What began as a small problem is spreading across the entire globe. | | Already, the shock has touched a single locality. | Already, the shock has touched the nation as a whole. | Already, the shock has touched every country on Earth. | | Consequences now extend across one local area. | Consequences now extend across the country at large. | Consequences now extend across the global community. | | The situation continues to escalate across a single town. | The situation continues to escalate across the whole country. | The situation continues to escalate across the entire world. | | Fallout from the event now reaches one neighborhood. | Fallout from the event now reaches the entire nation. | Fallout from the event now reaches every continent. | | Its effects are rippling across a single local community. | Its effects are rippling across the entire country. | Its effects are rippling across countries around the globe. | | A growing disturbance is being felt throughout one city. | A growing disturbance is being felt throughout the whole nation. | A growing disturbance is being felt throughout nations worldwide. | | The threat now spans a single district. | The threat now spans the country from coast to coast. | The threat now spans the whole planet. | | By every measure, the incident is affecting one village. | By every measure, the incident is affecting the country as a whole. | By every measure, the incident is affecting the entire globe. | | Upheaval is playing out across a single locality. | Upheaval is playing out across the nation as a whole. | Upheaval is playing out across every country on Earth. | | The disruption stretches across one local area. | The disruption stretches across the country at large. | The disruption stretches across the global community. | | There are serious consequences for a single town. | There are serious consequences for the whole country. | There are serious consequences for the entire world. | | Mounting strain is being felt across one neighborhood. | Mounting strain is being felt across the entire nation. | Mounting strain is being felt across every continent. | | The downturn has now reached a single local community. | The downturn has now reached the entire country. | The downturn has now reached countries around the globe. | | The crisis is reverberating across one city. | The crisis is reverberating across the whole nation. | The crisis is reverberating across nations worldwide. | ## Step 2: embed your geographic reference texts with OpenAI An embedding turns each sentence into a list of numbers that captures its meaning. We use OpenAI's `text-embedding-3-large` (3,072 dimensions) for a specific reason: it is frozen, with a September 2021 knowledge cutoff. Foreknowledge bias is the error of scoring a past event using information about how it actually turned out. Because this model's knowledge stops in September 2021, any event after that date is scored with no idea of what came next, so a backtest over recent history carries none of it. A live LLM, retrained on everything since, gives you no such guarantee. ## Step 3: compute cosine similarity to the geographic reference texts Cosine similarity measures how close two embeddings point in the same direction: high means similar, low means different. Score the input against all sixty references. The raw numbers are noisy and overlap heavily across buckets, so you cannot just read off an answer. ![Strip plot of the input sentence's 60 raw cosines, 20 per bucket, colored local, national, global. The clouds are noisy and overlap in the middle.](https://nosible.com/images/2026/06/walkthrough_step3_geo_raw.png) Strip plot of the input sentence's 60 raw cosines, 20 per bucket, colored local, national, global. The clouds are noisy and overlap in the middle. ## Step 4: apply the tanh trick to squash background noise Even two unrelated sentences will score around 0.3 just for sharing the same clipped, declarative register. Pass a tanh gate over the cosines, with a band from 0.25 to 0.50: anything below collapses toward 0, anything above saturates toward 1. The background noise is squashed and the real matches stand out. ![The same 60 cosines after gating. The background is squashed toward 0 and the strong global matches push toward 1.](https://nosible.com/images/2026/06/walkthrough_step4_geo_gated.png) The same 60 cosines after gating. The background is squashed toward 0 and the strong global matches push toward 1. ## Step 5: use a cubic mean to amplify the strongest signal A bucket is many sentences, so turn each bucket's gated cosines into one number with a cubic power mean (p=3), which leans the aggregate toward the strongest matches instead of letting weak ones drag it down. The input scores **local 0.315, national 0.417, global 0.739**. The argmax is **global**. A dead-zone floor would return a clean "none" if no bucket cleared it, so off-topic events are left unassigned rather than forced. ![Bar chart of the three bucket scores: local 0.315, national 0.417, global 0.739. The argmax is global.](https://nosible.com/images/2026/06/walkthrough_step5_geo_scores.png) Bar chart of the three bucket scores: local 0.315, national 0.417, global 0.739. The argmax is global. ## Step 6: write mirror texts for systemic and idiosyncratic labels The second job is a single axis with two poles, so write the two sets as **mirror images**: identical sentences except for the words that flip single-company to whole-market. These are the real production pairs. Topics vary across the pairs on purpose, so that averaging cancels the topic and leaves only the single-company versus whole-market flip. | Idiosyncratic (one company) | Systemic (whole market) | | --- | --- | | Analysts warn that a single company could default on its debt | Analysts warn that companies across the economy could default on their debt | | Investors caution that one firm's share price may collapse | Investors caution that the entire stock market may collapse | | Economists predict that a single business might be forced into bankruptcy | Economists predict that businesses nationwide might be forced into bankruptcy | | Analysts fear that one company may miss its profit forecast | Analysts fear that companies across every sector may miss their profit forecasts | | Officials forecast that a single bank may soon fail | Officials forecast that the whole banking system may soon fail | | Strategists believe that an individual company could soon see its credit rating cut | Strategists believe that issuers across the market could soon see their credit ratings cut | | Analysts flag that a single retailer could run out of cash | Analysts flag that retailers throughout the industry could run out of cash | | Investors fret that one firm's bonds may become worthless | Investors fret that corporate bonds across the market may become worthless | | Economists say a single manufacturer might halt production | Economists say manufacturers nationwide might halt production | | Analysts expect that a particular company may cut thousands of jobs | Analysts expect that companies across the economy may cut millions of jobs | | Officials warn that one firm could soon be hit with a record fine | Officials warn that firms across the industry could soon be hit with record fines | | Investors caution that a single stock could be wiped out | Investors caution that the broad equity market could be wiped out | | Analysts predict that one company's profits may evaporate | Analysts predict that corporate profits across the economy may evaporate | | Strategists fear that a single firm might lose its market access | Strategists fear that firms everywhere might lose their market access | | Economists forecast that a single borrower may be unable to refinance | Economists forecast that borrowers across the market may be unable to refinance | | Analysts believe that one company may soon face a crippling lawsuit | Analysts believe that companies across the sector may soon face crippling lawsuits | | Officials flag that a single firm could soon be caught in a fraud scandal | Officials flag that fraud scandals could soon spread across the whole industry | | Investors fret that one company could slash its dividend | Investors fret that companies across the market could slash their dividends | | Analysts say a single business may lose its biggest customer | Analysts say businesses throughout the economy may lose their biggest customers | | Strategists expect that an individual stock might plunge overnight | Strategists expect that the entire market might plunge overnight | ## Step 7: embed your mirror texts with OpenAI again Embed both sets the same way. If you project the raw embeddings down to two dimensions, the two sets sit on top of each other. Raw cosine cannot separate them, because both sets are written in the same financial register and that shared style dominates the vectors. ![A 2D projection of the raw mirror-set embeddings. The idiosyncratic and systemic clouds overlap heavily.](https://nosible.com/images/2026/06/walkthrough_step7_scope_raw_2d.png) A 2D projection of the raw mirror-set embeddings. The idiosyncratic and systemic clouds overlap heavily. ## Step 8: compute the systemic and idiosyncratic contrast vectors The fix is common-mode neutralization. Subtract the mean of all the reference vectors from each one, then renormalize. That removes the shared component the two sets carry. The shared coordinate drops from **+0.71 to +0.005**, and viewed along the contrast direction the method actually uses, the two sets now split cleanly. ![The same sets after neutralization, projected onto the contrast direction. Idiosyncratic on the right, systemic on the left, cleanly separated.](https://nosible.com/images/2026/06/walkthrough_step8_scope_neut_2d.png) The same sets after neutralization, projected onto the contrast direction. Idiosyncratic on the right, systemic on the left, cleanly separated. ## Step 9: compute the similarity to the contrast vectors Score the input as a signed cosine contrast: its mean cosine to the neutralized idiosyncratic set minus its mean cosine to the neutralized systemic set. Positive means idiosyncratic (one company), negative means systemic (the whole market). Our input scores **−0.068**, landing on the systemic side of zero, which is right: a synchronized shock across every major economy is the whole market, not one company. ![Signed-contrast number line. Our running event sits at −0.068, left of zero on the systemic side of an axis that runs from systemic (whole market) on the left to idiosyncratic (one company) on the right.](https://nosible.com/images/2026/06/walkthrough_step9_number_line.png) Signed-contrast number line. Our running event sits at −0.068, left of zero on the systemic side of an axis that runs from systemic (whole market) on the left to idiosyncratic (one company) on the right. ## Step 10: fuse the signals together to get the full picture The two reads are independent, so place each event on a plane: geographic scale on one axis, systemic versus idiosyncratic on the other. The input lands firmly in the global and systemic corner. ![A 2D scatter of the verification events plus the running input, each labelled by name and positioned by geographic scale and scope. The running input sits in the global, systemic corner.](https://nosible.com/images/2026/06/walkthrough_step10_fusion.png) A 2D scatter of the verification events plus the running input, each labelled by name and positioned by geographic scale and scope. The running input sits in the global, systemic corner. ## Verification Let's apply the method to 18 sentences to check that it works: three geographies times two scopes, three samples each. The national and global cases are real warnings pulled straight from the NOSIBLE World leaderboard; the local cases are authored, because the leaderboard has nothing at town scale. We score every one through the full pipeline and compare the prediction to a label assigned by hand. The last column shows what a language model ( `gemini-2.5-flash`, temperature 0) returns for the same sentence. | Sentence | Intended | Method | LLM (gemini-2.5-flash) | scope score | | --- | --- | --- | --- | --- | | Every employer in a small mill town warned of mass layoffs after the factory anchoring the local economy announced its closure. | local / systemic | local / idiosyncratic | local / systemic | +0.010 | | Shopkeepers across a small coastal town warned that a collapsed tourist season had pushed the whole local economy to the brink. | local / systemic | local / systemic | local / systemic | −0.057 | | Officials in a small farming town warned that a failed harvest would ripple through every business on the main street. | local / systemic | local / systemic | local / systemic | −0.029 | | The owner of a family bakery in a small town warned that surging rents would force the century-old shop to close. | local / idiosyncratic | local / idiosyncratic | local / idiosyncratic | +0.017 | | A single hardware store in a small town warned it would shut for good after a fraud drained its accounts. | local / idiosyncratic | local / idiosyncratic | local / idiosyncratic | +0.051 | | A local manufacturer in one town warned of layoffs after losing its only major contract. | local / idiosyncratic | local / idiosyncratic | local / idiosyncratic | +0.065 | | Analysts and consumer groups warn that persistent inflation and high energy costs will strain households and create fragility in Italy's housing market. | national / systemic | national / systemic | national / systemic | −0.027 | | Economic forecasters warn that UK unemployment will rise to a multi-year high as economic growth stalls in the coming months. | national / systemic | national / systemic | national / systemic | −0.054 | | Economists warn that Canada's economy will face sustained slower growth and a potential recession as artificial supports fade. | national / systemic | global / systemic | national / systemic | −0.059 | | Industry leaders warn that the liquidation of sugar producer Tongaat Hulett could trigger a failure of South Africa's sugar industry. | national / idiosyncratic | national / idiosyncratic | national / systemic | +0.022 | | Modella Capital warns that retailer TG Jones faces administration and collapse unless its lenders approve a restructuring plan. | national / idiosyncratic | local / idiosyncratic | national / idiosyncratic | +0.038 | | The Postmaster General warns that the United States Postal Service will run out of cash and be unable to pay its workers within a year. | national / idiosyncratic | national / systemic | national / systemic | −0.016 | | The IMF and a former central-bank governor warn that the global economy is unprepared for increasingly frequent and unpredictable shocks. | global / systemic | global / systemic | global / systemic | −0.031 | | Bond strategists warn that persistent inflation and geopolitical tension will keep government-bond yields elevated and push up borrowing costs worldwide. | global / systemic | global / systemic | global / systemic | −0.050 | | Economists warn that a synchronized slowdown will drag down corporate earnings across every major economy. | global / systemic | global / systemic | global / systemic | −0.058 | | Analysts warn that Tesla's stock price could plunge by more than 60% over the next year. | global / idiosyncratic | global / idiosyncratic | national / idiosyncratic | +0.043 | | Analysts warn that the highly anticipated IPO of SpaceX will struggle to outperform the market after its debut. | global / idiosyncratic | global / idiosyncratic | national / idiosyncratic | +0.028 | | Samsung's leadership and labour unions warn that planned strikes over pay will disrupt global semiconductor supply chains. | global / idiosyncratic | global / systemic | global / systemic | −0.005 | **Method: geography 16 of 18, scope 15 of 18, and both axes right at once on 13 of 18.** These are messy real-world warnings, not toy sentences, and the misses are the genuinely hard ones. The hardest are single companies whose trouble bleeds into a whole sector: Samsung's strike "disrupting global semiconductor supply chains" and the US Postal Service running out of cash both read as systemic, not idiosyncratic. On geography, Canada's macro warning tips global and the UK retailer TG Jones reads local rather than national. A current frontier model, Google's `gemini-2.5-flash`, lands in exactly the same place: **geography 16 of 18, scope 15 of 18, and both at once 13 of 18**, a dead heat with the geometry on all three. It trips on the same Samsung and Postal Service scope calls, also reads the Tongaat Hulett warning as systemic, and drops the two unmistakably global companies, Tesla and SpaceX, into the national bucket. A frozen embedding and a handful of cosines hold their own against the model they are meant to replace, while staying deterministic, auditable, and effectively free. ## Why not just ask an LLM? A language model can label an event too. For doing this at scale, on a corpus you have to stand behind, the geometry wins on five counts: - **Deterministic and reproducible.** The same sentence always yields the same numbers. An LLM's answer drifts between calls and model versions, and you cannot audit why it chose a label. - **Defensible.** Every output is a cosine, a mean, or a subtraction. You can re-derive any score by hand and explain it to a regulator or a client. - **Cheap and fast at scale.** One embedding per document, then linear algebra for every bucket and every axis. The LLM route is one call per document per question, which does not scale across millions of documents and many features. - **No training and no labelled data.** You write reference sentences. That is the entire setup. - **No foreknowledge bias.** The model's knowledge stops in September 2021, so when you backtest on events after that date it cannot have seen how they played out, and the features leak nothing from the future. ## Why this works Every number here is a cosine, a mean, or a subtraction. There is no training set to curate and nothing deciding the answer; the embedding is frozen, so the same text always yields the same features, reproducible and auditable by anyone. Two tricks cover the whole problem: a multiclass relevance score that sorts events into buckets (one set per bucket, gate, cubic mean, argmax), and a binary contrast that buckets the risk (two mirror sets, neutralize, signed cosine). Pick your problem, write the sentences, run the trick. --- *Every figure was generated from real `text-embedding-3-large` vectors using the operations above. The geographic reference texts are authored for this post; the verification set pairs authored local examples with real warnings from the NOSIBLE World leaderboard; the systemic and idiosyncratic mirror sets are taken straight from production.* [All Research](https://nosible.com/blog) Related Research ![Thirty semantic factors rendered as daily time series on a dark grid, from geopolitical risk to firm-level political risk.](https://nosible.com/images/2026/08/semantic-factors-hero.png) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ### [The Future That Could Have Been: Turning Web Text into Semantic Stock Betas](https://nosible.com/blog/the-future-that-could-have-been-turning-web-text-into-semantic-stock-betas) 2026-08-05 11 min read ![Daily NOSIBLE Trade Policy Uncertainty index compared with the published index](https://nosible.com/images/2026/06/nosible-tpu-daily.png) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ### [An Embedding-Based Approach to Trade and Economic Policy Uncertainty](https://nosible.com/blog/an-embedding-based-approach-to-trade-and-economic-policy-uncertainty) 2026-06-17 23 min read ![High-resolution chart of selected 35B-A3B serving records across a 350-plus experiment campaign, from Qwen 3.5 FP8 on H200 at $0.218400 per million completion tokens in Experiment 006 to Qwen 3.6 source NVFP4 at $0.125959 in Experiment 332](https://nosible.com/images/2026/08/qwen-35b-a3b-serving-campaign-hero.svg) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) ### [What 350+ Experiments Taught Us About Serving Qwen 3.6 35B-A3B](https://nosible.com/blog/what-350-experiments-taught-us-about-serving-qwen-3-6) 2026-08-17 15 min read > Turn OpenAI sentence embeddings into clean geographic and systemic-risk features with two deterministic, training-free scoring techniques. **URL:** https://nosible.com/blog/the-contrastive-geometry-of-risk --- --- title: "An Embedding-Based Approach to Trade and Economic Policy Uncertainty" description: "Rebuild the Fed’s Trade Policy Uncertainty index from 14.9 million NOSIBLE World events using embeddings, then extend the method to broader policy risk." url: "https://nosible.com/blog/an-embedding-based-approach-to-trade-and-economic-policy-uncertainty" --- [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) # An Embedding-Based Approach to Trade and Economic Policy Uncertainty [Matthew Dicks](https://www.linkedin.com/in/matthewdicks98/) 2026-06-17 23 min read Copy as Markdown ![Daily NOSIBLE Trade Policy Uncertainty index compared with the published index](https://nosible.com/images/2026/06/nosible-tpu-daily.png) In April 2025 the United States put tariffs on nearly every trading partner. One of the ways to measure that stress is the [Trade Policy Uncertainty index](https://www.matteoiacoviello.com/tpu.htm), built by Federal Reserve economists. That month it hit the highest value in its 65-year history. 11.5% of all newspaper articles were about trade policy uncertainty. The index counts words. An article counts if it contains a trade-policy word and an uncertainty word. It reads seven major newspapers, back to 1960. This post reconstructs that index from the NOSIBLE World event database using only embeddings. We score 14.9 million news events against five sentences, with no keyword lists and no language model reading each article. The result matches the published benchmark at 0.87 on monthly levels and 0.82 on monthly changes, against the 0.96 and 0.70 that the two official versions reach with each other. The same method, extended to sixty sentences, rebuilds the broader [Economic Policy Uncertainty index](https://academic.oup.com/qje/article/131/4/1593/2468873), matching the published US series at 0.77 on monthly levels. ## The benchmark [Caldara, Iacoviello, Molligo, Prestipino and Raffo (2020)](https://www.matteoiacoviello.com/tpu_files/TPU_PAPER.pdf), in the *Journal of Monetary Economics*, built the index. Its history lines up with the record, including the Nixon shock of 1971, the NAFTA talks, and the 2018-19 US-China trade war. A second team checked it. [Baker, Bloom and Davis](https://www.policyuncertainty.com/trade_uncertainty.html) (BBD) built a rival trade-policy uncertainty index from different newspapers with different word lists. Over our window the two indices agree at 0.96 on monthly levels and 0.70 on month-to-month changes. Both read mostly the same English-language newspapers with hand-tuned keyword lists. Our question is whether an embedding-based approach over the NOSIBLE event database can produce a similar signal. Everything in this post is measured on the window where all three indices exist, 2015 to 2026, on the same dates for every pair. ## The method Every event in NOSIBLE World carries an embedding, a list of numbers describing what the event means, computed once when the event is created. It is an OpenAI `text-embedding-3-large` vector, stored in the event's `oai_vector` field. The model returns 3,072 dimensions. We keep the first 1,024 for speed, which the model's Matryoshka training makes safe, and L2-normalize them so cosine similarities stay meaningful. One real-world event is one record, no matter how many outlets cover it, and each record stores how many separate publishers did. We call that breadth, the `total_netlocs` field. The whole index is five sentences, compared against each event's embedding. Three sentences define the topic. An event is about trade policy if its embedding is close to any of them. ``` tariffs: "Import tariffs and duties: a government imposing, raising, threatening or suspending tariffs, customs duties, import quotas, surcharges or fees on goods imported from other countries" agreements: "Trade negotiations, trade agreements and trade treaties between countries: free trade deals being negotiated, renegotiated, signed, ratified, suspended or abandoned, and trade talks between governments" disputes: "Trade disputes, trade wars and protectionism: retaliatory tariffs, anti-dumping measures, import barriers, export controls and restrictions, boycotts of foreign goods, and trade complaints before the WTO" ``` Two more sentences define the uncertainty axis. They replace the paper's "AND uncertainty words" rule. Each event is scored against both, and the gap between the two scores becomes a weight. ``` uncertain: "Trade policy uncertainty: the outlook for tariffs, trade agreements and trade rules is uncertain, unclear and unpredictable; threatened tariffs may or may not happen, trade negotiations are stalled or at risk of collapse, and businesses cannot plan for what trade policy comes next" certain: "Trade policy certainty and stability: a trade agreement is concluded and ratified, a trade dispute is resolved, tariffs are removed or finalized, and governments give businesses a clear, settled and predictable trade policy outlook" ``` Each day, the index is the share of publisher attention spent on trade-policy events, tilted toward the ones framed as uncertain. ``` relevant(e) = max cosine(event e, the 3 topic phrases) >= 0.35 polarity(e) = tanh((uncertain_sim - certain_sim) / 0.1) # -1 resolved .. +1 uncertain w_unc(e) = (1 + polarity(e)) / 2 # 0 .. 1 sum of breadth(e) * w_unc(e) over relevant events on day t NOSIBLE-TPU(t) = ------------------------------------------------------------- trailing 12-month average of total daily breadth ``` The denominator strips out the growth of the corpus itself, the same way our [geopolitical risk study](https://nosible.com/blog/rebuilding-the-geopolitical-risk-index-from-nosible-world) did. The 0.35 cutoff is not load-bearing. Moving it from 0.30 to 0.40 changes every correlation by about 0.02. That is the whole method. Scoring all 14.9 million events from 2015 to 2026 takes about twelve minutes on a laptop. The sentences work in every language the database holds, because the embeddings do. The trade index is scored over every event worldwide, not a US subset. The economic-policy index later in this post is the one exception, restricted to US events to match the published series. ## It matches the benchmark Start with the daily series. The published index dates each article by its print day. Ours dates each event by the day its coverage peaked. ![Daily NOSIBLE-TPU against the published daily Trade Policy Uncertainty index, both as 30-day moving averages rebased to a common mean, 2015 to 2026. The two lines move together and correlate at 0.86 over the full window. Both stay flat to 2017, rise on the 2018 steel and China tariffs and the 2019 escalations, go quiet through 2024, then spike to record highs in 2025, with Liberation Day the tallest point. The shapes match almost exactly. The main difference is that NOSIBLE runs higher through the 2018-19 trade war.](https://nosible.com/images/2026/06/nosible-tpu-daily.png) Daily NOSIBLE-TPU against the published daily Trade Policy Uncertainty index, both as 30-day moving averages rebased to a common mean, 2015 to 2026. The two lines move together and correlate at 0.86 over the full window. Both stay flat to 2017, rise on the 2018 steel and China tariffs and the 2019 escalations, go quiet through 2024, then spike to record highs in 2025, with Liberation Day the tallest point. The shapes match almost exactly. The main difference is that NOSIBLE runs higher through the 2018-19 trade war. The two daily lines correlate at 0.86 across the full window. Both sit flat through 2015 to 2017, then climb together on the 2018 steel and aluminium tariffs and the first China rounds. They peak again on the 2019 escalations, fall quiet through 2020 to 2024, and explode in 2025. Liberation Day in April 2025 is the single tallest spike in eleven years, followed by the Geneva truce and the 2026 Supreme Court ruling on the emergency-powers tariffs. The shapes match closely; only the height differs. NOSIBLE runs higher through the 2018-19 trade war, because that war was a larger share of a global, then-smaller corpus than of seven US and UK papers. Now the monthly view, with the second published measure added. ![Three monthly measures of trade policy uncertainty, rebased to a common mean and aligned on the same dates, the published TPU, the Baker-Bloom-Davis trade-policy EPU, and NOSIBLE-TPU. All three track each other and peak together in April 2025. Against the published TPU the two official measures agree at 0.96 on levels and 0.70 on changes, while NOSIBLE matches it at 0.87 and 0.82, so on month-to-month changes it tracks the benchmark more closely than the two official measures track each other.](https://nosible.com/images/2026/06/nosible-tpu-threeway.png) Three monthly measures of trade policy uncertainty, rebased to a common mean and aligned on the same dates, the published TPU, the Baker-Bloom-Davis trade-policy EPU, and NOSIBLE-TPU. All three track each other and peak together in April 2025. Against the published TPU the two official measures agree at 0.96 on levels and 0.70 on changes, while NOSIBLE matches it at 0.87 and 0.82, so on month-to-month changes it tracks the benchmark more closely than the two official measures track each other. The published TPU, the BBD trade-policy EPU, and NOSIBLE-TPU move together across the whole window. Against the published TPU, BBD scores 0.96 on levels and 0.70 on changes. NOSIBLE scores 0.87 and 0.82. On the month-to-month changes, our index agrees with the benchmark more closely than the two published measures agree with each other. All three rank April 2025 as their highest month. Every pair, on the same dates: | pair | monthly level / change | quarterly level / change | | --- | --- | --- | | TPU ~ BBD (the bar) | 0.96 / 0.70 | 0.98 / 0.93 | | **TPU ~ NOSIBLE** | **0.87 / 0.82** | **0.90 / 0.83** | | BBD ~ NOSIBLE | 0.79 / 0.52 | 0.85 / 0.71 | NOSIBLE is closer to TPU than to BBD on every change metric. ## It spikes in the same months A correlation can hide a lot, so we check the months one by one. All three indices rank April 2025 as their top month. Every one of the twelve episodes we flag lands above the 60th percentile in all three, and every pair of indices peaks together with no lag. Here is a selection, ranked out of 135 months. | episode | month | TPU rank | BBD rank | NOSIBLE rank | | --- | --- | --- | --- | --- | | Section 232 steel and aluminium | 2018-03 | 18 | 34 | 11 | | China Section 301, round one | 2018-06 | 22 | 28 | 4 | | August 2019 escalation | 2019-08 | 19 | 15 | 10 | | Phase One deal signed | 2020-01 | 38 | 30 | 53 | | IEEPA tariffs on Canada, Mexico, China | 2025-02 | 7 | 11 | 5 | | Liberation Day | 2025-04 | 1 | 1 | 1 | | Geneva truce | 2025-05 | 2 | 3 | 6 | | IEEPA Supreme Court hearing | 2025-11 | 14 | 7 | 44 | | IEEPA struck down | 2026-02 | 11 | 9 | 15 | Two rows stand out. NOSIBLE ranks the 2018-19 trade war higher, because the trade war was a bigger share of a global corpus than of seven US and UK papers. And it ranks the late-2025 litigation months lower, when trade-relevant coverage was lighter. ## A signed uncertainty axis The paper's key rule is the AND. Trade words alone do not count, an uncertainty word has to appear too. Our version of that rule is the per-event weight, built from the two opposing sentences. Unlike the original, it has a sign. The weight reflects the events, not the sentences. On the 0-to-1 weight, where 0.5 is neutral, events clearly not about trade average 0.44, just below neutral, while events that pass the trade filter average 0.63, tilted toward uncertainty. It also discriminates rather than pushing everything toward the middle. Among the events the index counts, 60% land clearly on the uncertain side of the weight, above 0.6, and 21% clearly on the settled side, below 0.4. The remaining fifth sit in the neutral band between, where the uncertain and certain sentences score about evenly. The chart below shows the net polarity of each day's trade coverage. Every relevant event gets a polarity score, from -1 fully resolved to +1 fully uncertain, and the daily reading is the breadth-weighted average over that day's events. ``` sum of breadth(e) * polarity(e) over relevant events on day t net_polarity(t) = --------------------------------------------------------------- sum of breadth(e) over relevant events on day t ``` The raw line is noisy, so we add a state-dependent smoother on top. When polarity jumps sharply toward uncertainty, 1.5 standard deviations above its trailing year with no look-ahead, the smoother switches to a fast 3-day half-life and picks up the shock at once. On calm days, and on moves toward resolution, it runs a slow 30-day half-life. The asymmetry is there to capture the spikes quickly. ![The daily net polarity of trade-policy coverage, scored from -1 when coverage is resolution-framed to +1 when it is uncertainty-framed, with a state-dependent smoothing overlay. The line runs positive during escalations, +0.59 at the 2018 China tariffs and +0.48 on Liberation Day, sits near zero through the quiet stretch of 2021, and turns negative when a dispute resolves, such as the January 2020 Phase One deal. This signed direction is what the published level-only index cannot show, telling a resolved story apart from a simply quiet news week.](https://nosible.com/images/2026/06/nosible-tpu-polarity.png) The daily net polarity of trade-policy coverage, scored from -1 when coverage is resolution-framed to +1 when it is uncertainty-framed, with a state-dependent smoothing overlay. The line runs positive during escalations, +0.59 at the 2018 China tariffs and +0.48 on Liberation Day, sits near zero through the quiet stretch of 2021, and turns negative when a dispute resolves, such as the January 2020 Phase One deal. This signed direction is what the published level-only index cannot show, telling a resolved story apart from a simply quiet news week. The polarity reads +0.59 during the China 301 escalation and +0.48 on Liberation Day. It sits near zero through the quiet stretch of 2021. It turns negative toward resolution when the Phase One deal is signed in January 2020. The published index also drops that month, but a keyword count drops the same way in any quiet month when trade simply leaves the news. A level on its own cannot tell a resolved dispute from a slow news week, but the sign can. A fall toward resolution means a story closed; a reading near zero means it went quiet. ## When uncertainty runs high, tariffs tend to follow A fair question about any news-based index is whether it tracks something that turns into policy, or only how loudly the press is talking. The chart plots the realized US tariff rate, customs duties divided by imports of goods from the quarterly BEA NIPA tables (via [DBnomics](https://db.nomics.world/) ), against all three news indices on the same quarters. ![The realized US tariff rate, quarterly, plotted against all three trade-policy-uncertainty indices, 2014 to 2026. The tariff rate edges up through the 2018-19 trade war, from 1.4% to 3%, then explodes from 2.6% to 12.8% in 2025, the largest move since the 1970s. The news indices lead the realized rate. Each index's level lines up with the next quarter's change in the tariff rate at 0.61 for TPU, 0.59 for BBD and 0.66 for NOSIBLE. Higher uncertainty tends to be followed by higher tariffs.](https://nosible.com/images/2026/06/nosible-tpu-tariffs.png) The realized US tariff rate, quarterly, plotted against all three trade-policy-uncertainty indices, 2014 to 2026. The tariff rate edges up through the 2018-19 trade war, from 1.4% to 3%, then explodes from 2.6% to 12.8% in 2025, the largest move since the 1970s. The news indices lead the realized rate. Each index's level lines up with the next quarter's change in the tariff rate at 0.61 for TPU, 0.59 for BBD and 0.66 for NOSIBLE. Higher uncertainty tends to be followed by higher tariffs. The higher the uncertainty, the more likely tariffs actually arrive, and the larger the move when they do. Through the 2018-19 threat war the indices rose and the realized rate crept up with them, from 1.4% to 3%. In 2025, when uncertainty reached its highest level on record, the rate exploded from 2.6% to 12.8%, the largest move since the Nixon era. The news indices lead the realized rate. Each index's level lines up with the following quarter's change in the tariff rate at 0.61 for TPU, 0.59 for BBD and 0.66 for NOSIBLE. Elevated trade policy uncertainty does not guarantee tariffs, but the higher it runs, the more often they follow. ## Where it misses Two things did not line up. **The 2018-19 size.** After rebasing, NOSIBLE's trade-war months run about twice as high as the published index. The timing matches; the magnitude reflects whose attention you are counting, a global corpus rather than seven US and UK papers. **Late-2025 litigation.** While the legal challenge to the 2025 emergency-powers tariffs worked through the courts, NOSIBLE eased back toward normal while the published indices stayed elevated. The litigation drew far less trade-relevant coverage than the active tariff actions earlier in the year, so NOSIBLE's trade signal was only modestly above normal in those months. The November 2025 oral arguments are the clearest case, TPU rank 14 against NOSIBLE rank 44; the February 2026 ruling brought NOSIBLE back in line, rank 15 against TPU's 11. ## The harder test: all of economic policy Trade is one topic. The real question is whether the recipe holds on something that covers all of economic policy at once. So we ran it against [Baker, Bloom and Davis's Economic Policy Uncertainty index](https://www.policyuncertainty.com/) ( *Quarterly Journal of Economics*, 2016). Their method counts articles that contain an economy word, a policy word and an uncertainty word across ten leading US newspapers. From that one index they also build category indices. We rebuild two, national security and healthcare. The yardstick comes the same way it did for trade. The paper's own two US versions, the 10-paper monthly index and the roughly 1,500-paper Newsbank daily index aggregated to monthly, are two measurements of one thing from different corpora. Over our window they agree at 0.92 on levels and 0.65 on changes. Scaling from one topic to all of economic policy needed two changes to the trade recipe. First, name the decisions. A single "economic policy uncertainty" sentence fails. The defining US events of April 2020, the Fed cutting rates to zero and the first shutdown orders, embed as concrete acts, not as the abstract idea of economic policy, so they score below any usable cutoff. The fix is a list of decision types, each written as a few concrete sentences. We use ten levers: monetary policy, taxes, spending, debt and shutdowns, entitlements, regulation, financial regulation, trade, major economic laws, and shutting down or reopening the economy. An event counts if its embedding is close to any of them. Second, pair the uncertainty sentences topic by topic. For trade, one uncertain-versus-certain pair was enough, because both sentences were full of trade words. Across many topics a single generic pair is lopsided. Uncertainty shares a vocabulary across topics ("may", "unclear", "stalled"), but resolution is always written as the topic's own concrete act ("signs the bill", "imposes the tariffs", "cuts rates"). Articles rarely say "policy is now certain". So every lever is a matched pair. One sentence frames the decision as proposed, threatened or undecided. The other frames the same decision as done and in force. Each event reads its polarity off its best-matching pair. The full set is sixty sentences, listed in the appendix. One pair shows the shape. ``` tariffs, uncertain: "The government is threatening or proposing to impose, raise, suspend or lift tariffs, but it is unclear what it will actually do." tariffs, certain: "The government has imposed, raised, suspended or lifted tariffs, and the tariff decision has taken effect." ``` The rest is the trade recipe unchanged, with one change to the denominator to match the paper's own normalization. ``` relevant(e) = max cosine(event e, all lever sentences) >= 0.35 polarity(e) = tanh((uncertain_sim - certain_sim) / 0.1) # within the best-matching pair w_unc(e) = (1 + polarity(e)) / 2 sum of breadth(e) * w_unc(e) over relevant US events in month t NOSIBLE-EPU(t) = ------------------------------------------------------------------- trailing 12-month average of total US-attributed monthly breadth ``` This index is US-only on both lines. An event is US-attributed if its tagged primary country is the United States. The numerator is US events that are economic-policy-relevant, and the denominator is all US events. We scale US attention by US attention, the way the published index scales US articles by US articles. The category indices add one rule. A national-security or healthcare event must also sit close enough (0.25) to that category's own sentence pairs, matching the paper's category definition. One lever earns a note. A keyword count caught COVID for free, because every closure article had an economy word, a policy word and an uncertainty word somewhere in its text. An event embeds as its core meaning, with no such incidental words, so the nine levers taken from the paper's pre-2020 categories could not see "government orders businesses to close". The tenth lever covers shutting down or reopening the economy. We wrote it in timeless terms, with nothing about any one pandemic, and we keep it as a separate line on the chart so the choice is visible. ![NOSIBLE-EPU against the published US Economic Policy Uncertainty index, monthly and rebased, with a nine-lever version that drops the shutdown lever shown as a diagnostic line. NOSIBLE tracks the published index at 0.77 on levels and 0.53 on changes. Both rank April 2025 the highest month of the period, and both place the COVID shock of spring 2020 near the top. The two NOSIBLE lines agree everywhere except 2020, which shows that the entire COVID gap comes down to the single shutdown lever.](https://nosible.com/images/2026/06/nosible-epu-us.png) NOSIBLE-EPU against the published US Economic Policy Uncertainty index, monthly and rebased, with a nine-lever version that drops the shutdown lever shown as a diagnostic line. NOSIBLE tracks the published index at 0.77 on levels and 0.53 on changes. Both rank April 2025 the highest month of the period, and both place the COVID shock of spring 2020 near the top. The two NOSIBLE lines agree everywhere except 2020, which shows that the entire COVID gap comes down to the single shutdown lever. NOSIBLE-EPU tracks the published US index at 0.77 on levels and 0.53 on changes. The two indices agree on where the big spikes are. Both rank April 2025 as the highest month of the eleven years, and both put the COVID shock of March and April 2020 and the 2025 tariff run in their top few. The cyan line drops the shutdown lever. The two NOSIBLE lines agree everywhere except 2020, where the gap is large and positive in March, April and May, then fades to nothing by mid-2021 and shows no real divergence anywhere else in eleven years. The whole COVID gap comes down to that one lever. Where NOSIBLE runs hotter than the keyword count is on clear-cut decisions like the February 2025 tariffs. Where it runs cooler is the back half of 2025, the condition-driven stretch discussed below. | target | monthly level / change | | --- | --- | | The bar (the paper's own two US variants) | 0.92 / 0.65 | | **US vs published US EPU** | **0.77 / 0.53** | | National security vs published categorical | 0.83 / 0.59 | | Healthcare vs published categorical | 0.73 / 0.44 | ![The national-security and healthcare categories of Economic Policy Uncertainty, NOSIBLE against the published categorical series, monthly. National security tracks the published series at 0.83 on levels and 0.59 on changes, with both ranking April 2025 first. Healthcare is looser at 0.73 and 0.44 but agrees on the periods that dominate the series, the COVID spring of 2020 and the 2025 health-policy upheaval. Both categories come from the same sentence set as the headline index, with no extra word lists.](https://nosible.com/images/2026/06/nosible-epu-categories.png) The national-security and healthcare categories of Economic Policy Uncertainty, NOSIBLE against the published categorical series, monthly. National security tracks the published series at 0.83 on levels and 0.59 on changes, with both ranking April 2025 first. Healthcare is looser at 0.73 and 0.44 but agrees on the periods that dominate the series, the COVID spring of 2020 and the 2025 health-policy upheaval. Both categories come from the same sentence set as the headline index, with no extra word lists. The two category indices spike in the right places. National security tracks the published series closely through the 2025-26 cluster, with both ranking April 2025 first, carrying heavy sanctions-and-retaliation coverage, and March 2025 and March 2026 close behind in both. It scores 0.83 on levels and 0.59 on changes. Healthcare scores 0.73 and 0.44, and agrees on the two periods that dominate the series, the COVID spring of 2020 and the 2025 health-policy upheaval, the Medicaid changes in the tax bill, the ACA-subsidy fight, and the leadership turnover at HHS. It also picks up the May 2017 attempt to repeal the Affordable Care Act, a clear legislative fight the keyword count ranks far down. Healthcare is the loosest of the four indices. The replication is real, but looser than trade, and the gap has one cause. It is the difference between counting words and measuring events. The published EPU counts coverage of conditions whenever it mentions policy. A record-unemployment article counts as long as it has an economy word, a government word and an uncertainty word somewhere in the text. Our index counts only events that embed as policy decisions. So months driven by conditions, like spring 2020's jobless-claims records or late 2025's background tariff mentions, sit below the published series. The category gaps come from the same difference. The published healthcare series counts the daily run of premium, Medicaid and insurance-cost coverage whenever it carries the right words, while NOSIBLE counts only health-policy decisions, so that ambient coverage does not lift our series the way it lifts theirs. Against those misses, the method scales in a way the keyword approach cannot. One sentence set and one cutoff build the headline index and both categories at once, while the paper hand-builds a separate word list for each category. The same index runs for any country as a group-by on the country field, in every language, while the paper needs local newspapers and a translated word list for each one. Adding a category or a country is nearly free. ## One recipe, two index families The two replications are the same recipe at two sizes. The Trade Policy Uncertainty index came from five sentences. The Economic Policy Uncertainty family, the headline index plus its national-security and healthcare categories, came from sixty. Neither used a keyword list or a language model reading each article. Both used only the embeddings the database already stores, compared against a handful of sentences. Both clear the bar we set, which is how closely the existing published indices match each other. NOSIBLE-TPU matches the published trade index as closely as the two published trade indices match each other, and more closely on the changes. NOSIBLE-EPU and its categories reach 80 to 90 percent of the agreement the published versions have among themselves, from one sentence set where the paper needs a word list per category. Both spike in the right months. Both carry a signed uncertainty-versus-resolution axis, a direction the published indices do not provide. And both run in every language at once. Nothing here is specific to trade or economic policy. The same recipe points at monetary-policy uncertainty, sanctions risk, energy-transition policy, or anything else people have measured with a keyword list. Writing the sentences takes an afternoon. The validation is the real work, and it now exists as a template. ## Future work The index is defined by two ingredients, the set of anchor sentences and the function that scores events against them. Both are deliberate but unoptimized, and each marks a direction for future work. The first is the anchor set. The trade-policy and economic-policy constructs were written by hand and calibrated against a single benchmark, so they are unlikely to be optimal. More precise definitions, and a broader set covering sub-topics the current phrases miss, should improve fidelity to the published series and narrow the documented gaps, in particular the shortfall on condition-driven rather than decision-driven coverage. Whether such anchors are best authored by hand, mined from the corpus, or learned against the benchmark is itself an open question. The second is the scoring function, which matters most for the signed certainty-versus-uncertainty axis. That axis maps the gap between two similarity scores through a fixed monotone transform, a modeling choice rather than a result, and a calibrated or learned mapping is the natural next step. Each of these directions can be evaluated the same way, against the published benchmarks. ## Work with us NOSIBLE turns the world's news into a structured, multilingual, de-duplicated event database, and this post used one slice of it. If you want access to that database, or a signal like these built for your own models, [start a trial](https://nosible.com/start-trial). You can explore the live data at [nosible.world](https://nosible.world/). ## References - Caldara, Dario, Matteo Iacoviello, Patrick Molligo, Andrea Prestipino and Andrea Raffo (2020). [The Economic Effects of Trade Policy Uncertainty](https://www.matteoiacoviello.com/tpu_files/TPU_PAPER.pdf). *Journal of Monetary Economics* 109, 38-59. Index and data: [matteoiacoviello.com/tpu](https://www.matteoiacoviello.com/tpu.htm). - Baker, Scott R., Nicholas Bloom and Steven J. Davis (2016). [Measuring Economic Policy Uncertainty](https://academic.oup.com/qje/article/131/4/1593/2468873). *Quarterly Journal of Economics* 131(4), 1593-1636. The headline US EPU, the daily US EPU, the categorical national-security and healthcare series, and the categorical trade-policy series used in the trade sections are all from [policyuncertainty.com](https://www.policyuncertainty.com/). - Caldara, Dario and Matteo Iacoviello (2022). Measuring Geopolitical Risk. *American Economic Review* 112(4). The sibling index, rebuilt from NOSIBLE World in [our previous post](https://nosible.com/blog/rebuilding-the-geopolitical-risk-index-from-nosible-world). - Customs duties and imports of goods: BEA NIPA tables via [DBnomics](https://db.nomics.world/). ## Appendix: the full EPU anchor set Everything you need to repeat the EPU results is here. Embed the sentences below with the same model used for the event embeddings (we use OpenAI text-embedding-3-large, truncated to its first 1,024 dimensions and re-normalized, which matches the vectors stored on every NOSIBLE World event). Score each event as the cosine against each sentence, then apply the formulas above. A lever's score is the highest of its two framings. Relevance is the highest score over all levers. Each event's polarity comes from the pair whose better framing scores highest. U is the uncertain sentence, C the certain one. To see this workflow as an editable research object, [explore 30 worked semantic factor definitions built from NOSIBLE World](https://nosible.com/semantic-factors). ``` MONETARY U: It is uncertain whether the central bank will raise, cut or hold interest rates, and markets do not know which way the decision will go. C: The central bank has announced its decision to raise, cut or hold interest rates, and markets know exactly where rates stand. U: The central bank may change its quantitative easing, balance-sheet or money-supply policy, but its plans remain unclear. C: The central bank has finalized its quantitative easing, balance-sheet and money-supply policy, and its plans are clear. U: A dispute over the central bank's independence or leadership has made the future of monetary policy unpredictable. C: The central bank's independence and leadership are secure, and the future of monetary policy is predictable. TAXES U: The government may cut taxes, raise taxes or change tax credits, and proposals are on the table, but nothing has been decided or enacted. C: The government has enacted its plan to cut taxes, raise taxes or change tax credits, and the new tax rules are law. U: A major overhaul of the tax code has been proposed and is being debated, and it is uncertain whether the tax reform will pass. C: A major overhaul of the tax code has been signed into law, and the tax reform is settled. SPENDING U: Proposals to raise, cut or freeze public spending and funding are on the table, but it is uncertain what the government will decide. C: The government has finalized its decision to raise, cut or freeze public spending and funding. U: The federal budget, appropriations or economic stimulus package has been proposed but is stuck in negotiations and may not pass. C: The federal budget, appropriations or economic stimulus package has been passed and signed into law. DEBT U: A debt-ceiling standoff or looming government shutdown has left government funding in doubt. C: A debt-ceiling deal has been reached and the government shutdown has been averted or ended, securing government funding. U: There is growing concern that the budget deficit, the national debt or the currency could spiral into a sovereign-debt crisis. C: The budget deficit and national debt are under control, and the sovereign-debt and currency pressures have been resolved. ENTITLEMENTS U: Proposed changes, cuts or reforms to public pensions, retirement benefits and welfare programs may or may not go ahead. C: Changes, cuts or reforms to public pensions, retirement benefits and welfare programs have been enacted and are final. REGULATION U: The government has proposed new rules or a rollback and is weighing whether to tighten regulation on business or to deregulate, and the outcome is uncertain. C: The government has enacted its decision to tighten regulation on business or to deregulate and roll back rules. U: An antitrust or competition case against companies is under way and its outcome is uncertain. C: The antitrust or competition case against the companies has been decided, and the enforcement outcome is final. FINANCIAL REGULATION U: Regulators have proposed new supervision and capital rules for banks and financial markets, but nothing is final and the rules remain in flux. C: Regulators have finalized the supervision and capital rules for banks and financial markets, and the rules are fixed. U: A bank bailout or rescue of the financial system is being debated and may or may not happen. C: The bank bailout and rescue of the financial system have been completed, and the financial system has been stabilized. TRADE U: The government is threatening or proposing to impose, raise, suspend or lift tariffs, but it is unclear what it will actually do. C: The government has imposed, raised, suspended or lifted tariffs, and the tariff decision has taken effect. U: A trade agreement between governments is under negotiation and may or may not be concluded. C: A trade agreement between governments has been signed and concluded. U: A trade war, trade dispute, sanctions or export controls could escalate, and nobody knows how far it will go. C: The trade war and trade dispute have been settled, and the sanctions and export controls are fixed and final. MAJOR ECONOMIC POLICY ACTIONS U: A major economic law or executive order on the economy has been proposed and is being considered, and may or may not be enacted. C: A major economic law or executive order on the economy has been enacted and is in force. U: A major economic reform or government intervention in the economy has been proposed and is under discussion, and its fate is uncertain. C: A major economic reform or government intervention in the economy has been carried out as planned. U: The leadership of an economic agency may be fired, replaced or shaken up, and the agency's direction is unclear. C: New leadership of the economic agency has been appointed and confirmed, and the agency's direction is settled. SHUTTING DOWN THE ECONOMY U: The government may shut down the economy and order businesses to close, and how long the shutdown of the economy will last is unclear. C: The government has shut down the economy and ordered businesses to close, and the shutdown is now in force. U: It is unclear when the government will reopen the economy and let businesses and economic activity resume. C: The government has reopened the economy, and businesses and economic activity are resuming as planned. CATEGORY: NATIONAL SECURITY U: A war, armed conflict or military invasion could break out or escalate, and it is unclear how the conflict will unfold. C: The war, armed conflict or military invasion has ended, and a ceasefire or peace agreement is in place. U: There are fears of terrorism and warnings that a terrorist attack may happen. C: The terrorist plot has been foiled and the attackers caught, and the terrorism threat has been contained. U: It is uncertain whether defense spending, the military budget or weapons procurement will be increased or cut. C: The defense budget has been approved, settling defense spending, the military budget and weapons procurement. U: Sanctions, blockades or embargoes are being threatened or considered, and it is unclear whether they will be imposed. C: The sanctions, blockades or embargoes have been imposed or lifted, and the decision is final. CATEGORY: HEALTHCARE U: The future of government health insurance and public healthcare programs is uncertain, with changes proposed but not decided. C: The changes to government health insurance and public healthcare programs have been enacted and finalized. U: Health insurance coverage, premiums, subsidies and healthcare costs may change, and people face an unpredictable outlook. C: Health insurance coverage, premiums, subsidies and healthcare costs have been set, and people have a clear outlook. U: Decisions on drug pricing, medical products and pharmaceutical or vaccine regulation are pending, leaving the industry uncertain. C: Decisions on drug pricing, medical products and pharmaceutical or vaccine regulation have been finalized and announced. U: The government may fire health officials, cut the funding of its health agencies or shake up their leadership, and the direction of health policy is unclear. C: The government has appointed and confirmed the leadership of its health agencies and secured their funding, and the direction of health policy is settled. ``` All correlations are computed on the common window, 2015-03 to 2026-05 monthly, on jointly non-null rows after rebasing each series to its own mean. The published US EPU is the news-based Fig-1 series from policyuncertainty.com, and the categorical series are the "National security" and "Health care" columns of their categorical file. The relevance threshold of 0.35 was chosen by a sweep against the published US index and applied unchanged to both categories. The category threshold is 0.25. [All Research](https://nosible.com/blog) Related Research ![Thirty semantic factors rendered as daily time series on a dark grid, from geopolitical risk to firm-level political risk.](https://nosible.com/images/2026/08/semantic-factors-hero.png) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ### [The Future That Could Have Been: Turning Web Text into Semantic Stock Betas](https://nosible.com/blog/the-future-that-could-have-been-turning-web-text-into-semantic-stock-betas) 2026-08-05 11 min read ![Signed contrast number line separating systemic and idiosyncratic risk](https://nosible.com/images/2026/06/walkthrough_step9_number_line.png) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ### [Two Tricks for Turning Sentence Embeddings into Clean Features](https://nosible.com/blog/the-contrastive-geometry-of-risk) 2026-06-18 14 min read ![High-resolution chart of selected 35B-A3B serving records across a 350-plus experiment campaign, from Qwen 3.5 FP8 on H200 at $0.218400 per million completion tokens in Experiment 006 to Qwen 3.6 source NVFP4 at $0.125959 in Experiment 332](https://nosible.com/images/2026/08/qwen-35b-a3b-serving-campaign-hero.svg) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) ### [What 350+ Experiments Taught Us About Serving Qwen 3.6 35B-A3B](https://nosible.com/blog/what-350-experiments-taught-us-about-serving-qwen-3-6) 2026-08-17 15 min read > Rebuild the Fed’s Trade Policy Uncertainty index from 14.9 million NOSIBLE World events using embeddings, then extend the method to broader policy risk. **URL:** https://nosible.com/blog/an-embedding-based-approach-to-trade-and-economic-policy-uncertainty --- --- title: "Trading Signals" description: "Research on operational signals and strategies derived from NOSIBLE data and public-web evidence." url: "https://nosible.com/blog/tag/trading-signals" --- /res Research # Trading Signals Research on operational signals and strategies derived from NOSIBLE data and public-web evidence. [All](https://nosible.com/blog) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Financial Sentiment](https://nosible.com/blog/tag/financial-sentiment) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [Trading Signals](https://nosible.com/blog/tag/trading-signals) [Web Search](https://nosible.com/blog/tag/web-search) 2 articles ![S&P 500 news-stress overlay with drawdown and signal thresholds](https://nosible.com/images/2026/06/news-stress-overlay-hero.png) [Trading Signals](https://nosible.com/blog/tag/trading-signals) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ## [Turning News into a Risk-On/Risk-Off Equity Signal](https://nosible.com/blog/turning-news-into-a-risk-on-risk-off-equity-signal) We built a risk-on/risk-off trading signal from the NOSIBLE event database that measures how much of the global news flow is about market-stress themes, holding equities when that reading is low and moving to T-bills when it spikes. Selected on 2010 to 2013 and tested on an untouched 2015 to 2026 window, it held the S&P 500's buy-and-hold return (+254% versus +269%) while cutting the maximum drawdown from −34% to −18% and raising the Sharpe ratio from 0.64 to 0.89. The same rule transfers unchanged to the Nasdaq and the Russell 2000. 2026-06-16 9 min read ![Magnifying-glass illustration representing vector search across company news](https://nosible.com/blog/illustrations/inspect.png) [Web Search](https://nosible.com/blog/tag/web-search) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) [Trading Signals](https://nosible.com/blog/tag/trading-signals) ## [Using Vector Search to See Signals in Company News](https://nosible.com/blog/using-vector-search-to-see-signals-in-company-news) How we use vector search to extract investment signals from a multi-terabyte company news dataset that currently contains over 55 million embeddings, 150+ million sentences, 4+ billion words, and 5+ billion GPT tokens. 2024-01-21 21 min read [All Research](https://nosible.com/blog) > Research on operational signals and strategies derived from NOSIBLE data and public-web evidence. **URL:** https://nosible.com/blog/tag/trading-signals --- --- title: "Turning News into a Risk-On/Risk-Off Equity Signal" description: "Build a point-in-time market-stress signal from NOSIBLE World that cuts S&P 500 drawdown and transfers unchanged to the Nasdaq and Russell 2000." url: "https://nosible.com/blog/turning-news-into-a-risk-on-risk-off-equity-signal" --- [Trading Signals](https://nosible.com/blog/tag/trading-signals) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) # Turning News into a Risk-On/Risk-Off Equity Signal [Gareth Warburton](https://www.linkedin.com/in/garethwarburton/) 2026-06-16 9 min read Copy as Markdown ![S&P 500 news-stress overlay with drawdown and signal thresholds](https://nosible.com/images/2026/06/news-stress-overlay-hero.png) We built a risk-on/risk-off trading signal from the NOSIBLE event database. Each day it measures how much of the global news flow is about market-stress themes, holds equities when that reading is low, and moves to T-bills when it spikes. We selected the parameters on 2010 to 2013 and tested on an untouched 2015 to 2026 window. On the S&P 500 the signal held buy-and-hold's total return (+254% versus +269%) while cutting the maximum drawdown from −34% to −18%, which raised the Sharpe ratio from 0.64 to 0.89. Five years into that test it stepped out of equities ahead of the 2020 COVID crash and back in as it passed, and the same rule transfers unchanged to the Nasdaq and the Russell 2000. The point is risk, not return: roughly the same long-run return with a much smaller drawdown. Everything below is enough to rebuild the signal, the rule, and the backtest from this post and the NOSIBLE event database alone. ## The signal: a daily stress reading from the news The signal uses three fields the database already stores on every event: its embedding ( `oai_vector`, from OpenAI `text-embedding-3-large` ), the number of distinct publishers that covered it ( `total_netlocs` ), and its date. Events are de-duplicated, so one real event is one record however many outlets run it. We define market stress as 17 short concepts and embed each once with the same model. The phrases are entity- and time-agnostic (no specific countries, crises, or dates) so no hindsight enters the backtest. The stored vectors are the full 3,072-dimensional `text-embedding-3-large` embedding, but the model is trained with Matryoshka representation learning, so the first 1024 dimensions are themselves a complete embedding. We use those: truncate both the event vectors and the anchors to the first 1024 dimensions, then L2-renormalise. The 17 anchors, verbatim: ``` credit_spreads Widening corporate credit spreads, surging default risk and stress in corporate bond markets liquidity A liquidity crisis or funding stress freezing financial markets deleveraging Forced selling, margin calls, deleveraging and liquidations cascading through markets recession Rising fears of recession, sharp economic slowdown or collapsing growth banking_credit A banking crisis, credit crunch, defaults or financial-system instability volatility_fear A surge in financial market volatility, fear, panic and investor anxiety equity_selloff A sharp sell-off, crash or plunge in global stock markets flight_to_safety A flight to safety and risk-off panic: investors dumping risky assets for government bonds, gold and reserve currencies sovereign_fx A sovereign-debt default or currency crisis destabilising markets monetary_shock An unexpected hawkish central bank or interest-rate shock tightening financial conditions rates_turmoil Turmoil in government bond markets: a bond rout, spiking yields and surging interest-rate volatility tail_hedging Investors rushing to buy downside protection and hedge against a market crash; spiking demand for portfolio insurance earnings_distress A wave of corporate profit warnings, earnings misses, bankruptcies and corporate distress employment_shock A labour-market shock: surging unemployment, mass layoffs or a sharply weakening jobs report tariffs New import tariffs, escalating protectionism and a trade war between major economies natural_disaster A major natural disaster, extreme-weather catastrophe or critical-infrastructure failure disrupting the economy war_conflict The outbreak or escalation of war, military conflict or a major geopolitical security crisis ``` For an event `e`, `relevance(e)` is the highest cosine similarity between its stored embedding and any of the 17 anchors, and `breadth(e)` is its `total_netlocs`. An event counts toward the day's reading only if `relevance(e) >= 0.30`. The day's raw stress reading is the breadth-weighted **share** of attention spent on stressful news: ``` sum over events on day t with relevance(e) >= 0.30 of relevance(e) * breadth(e) intensity(t) = ------------------------------------------------------------------------------------- sum over all events on day t of breadth(e) ``` Two choices matter. Breadth is the weight, so a story carried by 200 outlets counts once, weighted by 200, not as 200 events. And the reading is a share, not a count, so it does not drift up as the corpus grows: a count rises on crawl volume alone, a share of attention does not. ![Histogram of stress-relevance scores for real news events: the x-axis is each event's maximum cosine similarity to the 17 stress anchors, the green bars pile up near zero, and an amber dashed line marks the 0.30 relevance floor above which an event counts toward the daily reading](https://nosible.com/images/2026/06/news-stress-relevance-histogram.png) Histogram of stress-relevance scores for real news events: the x-axis is each event's maximum cosine similarity to the 17 stress anchors, the green bars pile up near zero, and an amber dashed line marks the 0.30 relevance floor above which an event counts toward the daily reading Relevance scores pile up near zero: most news is not about market stress. Only events past the `0.30` floor count. Two real headlines that clear it: | Stress relevance | Nearest anchor | Real headline | | --- | --- | --- | | 0.57 | `tariffs` | Trump Tariff Threats on Greenland Spark Global Trade War Fears | | 0.45 | `employment_shock` | US Jobless Claims Rise Marginally While Q3 Productivity Surges | ## Calibration: from a raw share to a causal score The raw share is not directly tradeable: its level is not comparable across years, and it is noisy. We turn it into a unitless, strictly-trailing z-score in three causal steps, none of which use future data: ``` 1. s(t) = 7-day trailing mean of intensity(t) (settle daily noise) 2. z(t) = (s(t) - median) / (1.4826 * MAD) (median and MAD over the trailing 252 trading days; 1.4826 puts it in sigma units) 3. z(t) = EWMA(z, span = 7) (so the position does not flicker) ``` ![The calibration pipeline in four stacked panels, top to bottom: the raw daily attention share, a 7-day rolling mean, the trailing 252-day robust z-score, and the EWMA-smoothed z that is actually traded, with the de-risk threshold drawn on the lower panels](https://nosible.com/images/2026/06/news-stress-calibration.png) The calibration pipeline in four stacked panels, top to bottom: the raw daily attention share, a 7-day rolling mean, the trailing 252-day robust z-score, and the EWMA-smoothed z that is actually traded, with the de-risk threshold drawn on the lower panels Top to bottom, the noisy raw share becomes a smoother, unitless score that spikes in 2020 and 2022 and is quiet between. The robust z rescales by the trailing window, so a reading of "2" means the same thing in a calm year and a volatile one. ## The rule A two-state machine on the calibrated z. The signal is lagged two trading days before it can trigger, and positions execute on the next bar, so every trade uses information knowable before it is placed: ``` state = LONG (fully in equities) if state == LONG and z > 1.75: -> OUT (move to BIL.US), days_out = 0 if state == OUT: days_out += 1 if days_out >= 5 and z < 0.25: -> LONG cost = 1 bp charged on |w(t) - w(t-1)| each time the position changes ``` When out, the defensive leg is `BIL.US` (SPDR 1-3 Month T-Bill ETF): the strategy earns the equity return when long and the T-bill return when out. That is the whole rule. The thresholds (exit 1.75, enter 0.25, min hold 5) were not hand-picked; the next section shows how they were selected and frozen. ![The 2025 tariff shock in two stacked panels: on top the S&P 500 growth of $1 with the out-of-equities window shaded red, and below the lagged news-stress z crossing the exit line (z greater than 1.75, dashed) in February and falling back under the re-enter line (z less than 0.25, dotted) in late May](https://nosible.com/images/2026/06/news-stress-episode-tariff-shock-2025.png) The 2025 tariff shock in two stacked panels: on top the S&P 500 growth of $1 with the out-of-equities window shaded red, and below the lagged news-stress z crossing the exit line (z greater than 1.75, dashed) in February and falling back under the re-enter line (z less than 0.25, dotted) in late May The 2025 tariff shock is one episode of the frozen rule on recent data. The stress z crossed the exit line in early February, the overlay moved to T-bills (shaded), sat out the roughly 19% peak-to-trough drop to the April low, and re-entered in late May once the z fell back under the re-enter line. It gave up part of the rebound. That is the cost. ## How the thresholds were chosen The rule is selected on a 2010 to 2013 training window, separated from the test window by a one-year (252 trading day) embargo so no test-period calibration window overlaps training data, then frozen. Every result below is on the untouched 2015-01-02 to 2026-06-01 test period, which selection never saw. On the train window only, we sweep a small grid and pick the cell that most improves the **Sortino ratio over buy-and-hold** (improvement over buy-and-hold, not the raw level, so the choice rewards timing rather than equity exposure): ``` exit_z in {1.75, 2.0, 2.25, 2.5, 2.75, 3.0} enter_z in {0.0, 0.25, 0.5, 0.75} min_hold in {3, 5, 10, 15} days ``` Selection is by plateau, not peak: take the top 20% of cells by train score and set each parameter to the grid value nearest the median of that set. This picks the robust centre of the good region rather than a single in-sample spike. The frozen result is exit `1.75` / enter `0.25` / hold `5`, with the news lagged 2 days, `BIL.US` as the defensive leg, and 1 bp per unit of turnover. ## A sanity check against the VIX Before any backtest: does a text-only stress reading line up with a market-based stress gauge at all? We plot the calibrated z against the VIX as a check that the signal is measuring market stress rather than noise. The VIX is not part of the strategy. ![The calibrated news-stress z (green, left axis) plotted against the VIX (amber, right axis) from 2010 to 2026, with the periods where the z sits above the de-risk threshold shaded red](https://nosible.com/images/2026/06/news-stress-signal-vs-vix.png) The calibrated news-stress z (green, left axis) plotted against the VIX (amber, right axis) from 2010 to 2026, with the periods where the z sits above the de-risk threshold shaded red The news-stress z (green) and the VIX (amber) rise and fall together through the major episodes (2011, 2015, 2020, 2022), even though the green line is built only from what the news is about and contains no price data. The check passes. Whether it helps a portfolio is the backtest below. ## Results: S&P 500 Over the 2015 to 2026 test period the overlay sits out the high-stress windows. It holds buy-and-hold's total return and cuts the maximum drawdown by nearly half, which raises the Sharpe ratio. ![S&P 500 news-stress overlay in three stacked panels, 2015 to 2026: log growth of $1 for buy-and-hold (white) versus the overlay (green), the drawdown path of each, and the lagged news-stress z with its exit and re-enter thresholds; red shading marks the days the overlay spent out of equities](https://nosible.com/images/2026/06/news-stress-overlay-sp500.png) S&P 500 news-stress overlay in three stacked panels, 2015 to 2026: log growth of $1 for buy-and-hold (white) versus the overlay (green), the drawdown path of each, and the lagged news-stress z with its exit and re-enter thresholds; red shading marks the days the overlay spent out of equities | Strategy | Sharpe | Sortino | Total return | Max drawdown | | --- | --- | --- | --- | --- | | Buy & Hold | +0.64 | +0.77 | +269% | −34% | | News-stress overlay | +0.89 | +1.01 | +254% | −18% | The overlay is out of equities about a quarter of the time, across 27 short episodes, with turnover under five round-trips a year. It is cheap to run and low-turnover. Same return, roughly half the drawdown. ## It transfers: Nasdaq and Russell 2000 The test of a news-based risk signal is whether it reads something about equity risk broadly or just one index. We take the rule frozen on the S&P 500 and apply it, unchanged, to two other markets. The signal is identical across all three; only the equity leg changes. ![Nasdaq Composite news-stress overlay versus buy-and-hold, 2015 to 2026: log growth of $1 (overlay in green, buy-and-hold in white), the drawdown of each, and the lagged news-stress signal, with the days out of equities shaded red](https://nosible.com/images/2026/06/news-stress-overlay-nasdaq.png) Nasdaq Composite news-stress overlay versus buy-and-hold, 2015 to 2026: log growth of $1 (overlay in green, buy-and-hold in white), the drawdown of each, and the lagged news-stress signal, with the days out of equities shaded red ![Russell 2000 (IWM) news-stress overlay versus buy-and-hold, 2015 to 2026: log growth of $1 (overlay in green, buy-and-hold in white), the drawdown of each, and the lagged news-stress signal, with the days out of equities shaded red](https://nosible.com/images/2026/06/news-stress-overlay-russell.png) Russell 2000 (IWM) news-stress overlay versus buy-and-hold, 2015 to 2026: log growth of $1 (overlay in green, buy-and-hold in white), the drawdown of each, and the lagged news-stress signal, with the days out of equities shaded red | Market | Sharpe (B&H → overlay) | Max drawdown (B&H → overlay) | Total return (B&H → overlay) | | --- | --- | --- | --- | | **S&P 500** (GSPC) | +0.64 → +0.89 | −34% → −18% | +269% → +254% | | **Nasdaq Composite** (IXIC) | +0.72 → +0.94 | −36% → −19% | +479% → +440% | | **Russell 2000** (IWM) | +0.41 → +0.67 | −41% → −22% | +186% → +254% | The Nasdaq matches the S&P result: near-flat total return, drawdown almost halved. The Russell 2000 is the stronger case: small caps fall harder in the flagged windows, so sitting them out cut the drawdown from −41% to −22% and raised total return from +186% to +254%. In all three markets the cut shows up in both halves of the test period, not just one, which is the opposite of what an overfit rule decaying out-of-sample would do. ## Honest accounting It reduces drawdown. It will not beat a strong bull market, and a sell-off that recovers costs return; that cost shows up in individual episodes. The test is honest by construction. The signal is strictly trailing, the news is lagged two trading days, and trades execute on the next bar. The rule was selected only on 2010 to 2013, frozen behind a one-year embargo, and measured solely on the untouched 2015 to 2026 window. ## One method, many signals The recipe generalises: define what you care about as a handful of phrases, score the world's already-embedded news against them, weight by how broadly each event was covered, and calibrate to a causal score. Point the same machinery at different anchors and you get a geopolitical-risk reading (which we have [matched against the Federal Reserve benchmark](https://nosible.com/blog/rebuilding-the-geopolitical-risk-index-from-nosible-world) ), commodity-supply stress, or a sector- or single-name news-pressure signal as an input to your own models. Each runs over the same database, across every language, with no language model reading each article. For editable examples of this workflow, [explore 30 worked semantic factor definitions built from NOSIBLE World](https://nosible.com/semantic-factors). ## Data and how to reproduce it Everything here is rebuildable from two sources: - **NOSIBLE event database** for the signal. Per de-duplicated event we use the stored embedding (`oai_vector`, OpenAI `text-embedding-3-large`), publisher breadth (`total_netlocs`), and event date, over the full daily history. Embed the 17 anchors above with the same model, truncate both the event vectors and the anchors to the first 1024 dimensions of the 3,072-dimensional Matryoshka embedding and L2-normalise, then compute `intensity(t)` and calibrate as specified. - **EODHD** end-of-day adjusted close for prices and the defensive leg: `GSPC.INDX` (S&P 500), `IXIC.INDX` (Nasdaq Composite), `IWM.US` (iShares Russell 2000), `BIL.US` (SPDR 1-3 Month T-Bill ETF), and `VIX.INDX` for the sanity-check figure. Returns are daily log returns; the overlay earns the equity leg when long and the T-bill leg when out, with 1 bp charged per unit of turnover. The full parameter set is fixed: relevance floor 0.30, rolling-mean window 7, robust-z window 252, EWMA span 7, news lag 2 trading days, exit z 1.75, enter z 0.25, min hold 5 days, 1 bp/turn, defensive leg `BIL.US`. ## Work with Nosible NOSIBLE turns the world's news into a structured, multilingual, de-duplicated event database, and this post used one slice of it. If you want access to that database, or a signal like this built for your own models, [start a trial](https://nosible.com/start-trial). You can explore the live data at [nosible.world](https://nosible.world/). [All Research](https://nosible.com/blog) Related Research ![Thirty semantic factors rendered as daily time series on a dark grid, from geopolitical risk to firm-level political risk.](https://nosible.com/images/2026/08/semantic-factors-hero.png) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ### [The Future That Could Have Been: Turning Web Text into Semantic Stock Betas](https://nosible.com/blog/the-future-that-could-have-been-turning-web-text-into-semantic-stock-betas) 2026-08-05 11 min read ![Signed contrast number line separating systemic and idiosyncratic risk](https://nosible.com/images/2026/06/walkthrough_step9_number_line.png) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ### [Two Tricks for Turning Sentence Embeddings into Clean Features](https://nosible.com/blog/the-contrastive-geometry-of-risk) 2026-06-18 14 min read ![Daily NOSIBLE Trade Policy Uncertainty index compared with the published index](https://nosible.com/images/2026/06/nosible-tpu-daily.png) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ### [An Embedding-Based Approach to Trade and Economic Policy Uncertainty](https://nosible.com/blog/an-embedding-based-approach-to-trade-and-economic-policy-uncertainty) 2026-06-17 23 min read > Build a point-in-time market-stress signal from NOSIBLE World that cuts S&P 500 drawdown and transfers unchanged to the Nasdaq and Russell 2000. **URL:** https://nosible.com/blog/turning-news-into-a-risk-on-risk-off-equity-signal --- --- title: "We Rebuilt the Geopolitical Risk Index with Nosible World" description: "Rebuild the geopolitical risk index from 13.2 million NOSIBLE World events and reproduce its global, country, country-pair and oil-risk signals." url: "https://nosible.com/blog/rebuilding-the-geopolitical-risk-index-from-nosible-world" --- [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) # We Rebuilt the Geopolitical Risk Index with Nosible World [Matthew Dicks](https://www.linkedin.com/in/matthewdicks98/) 2026-06-06 15 min read Copy as Markdown ![NOSIBLE geopolitical risk signal compared with published geopolitical risk indices](https://nosible.com/images/2026/06/nosible-gpr-vs-published.png) Markets move on geopolitics. A war, a coup, or a blockade moves oil, equities, and currencies within hours. But risk models cannot read the news. They take numbers, not headlines. So "the world feels dangerous right now" never reaches the model. We fixed that. We turned the Nosible World database into a single geopolitical risk signal, and it matches the benchmark that Federal Reserve economists built. It also matches the newer, harder version of that benchmark: the same risk scores, broken out for every country, every pair of countries, and for oil supply. We get there from data we already hold, in every language, without running a language model over every article. ## The problem: risk you can feel but cannot measure Every investor and policymaker knows geopolitics moves markets. The hard part is that geopolitical risk is a feeling, not a figure. There is a number for inflation, a number for unemployment, and a number for market volatility. There was never a number for how dangerous the world looks this month. Without one, the single force everyone agrees matters could not enter a forecast, a risk model, or a backtest. In 2018, two Federal Reserve economists, Dario Caldara and Matteo Iacoviello, [built that number](https://www.federalreserve.gov/econres/ifdp/files/ifdp1222.pdf). Their [Geopolitical Risk index](https://www.matteoiacoviello.com/gpr.htm) reads the daily newspaper record, measures how much of it is about wars, threats, and terrorism, and turns that into one series you can chart, compare, and test. For the first time, geopolitical risk became a variable you could put in a model. It turned out to matter. In the authors' own work, a rise in the index comes before lower business investment and hiring, weaker stock returns, and capital moving out of emerging markets toward safer ground. That is why it spread beyond academia and into the institutions that price risk for a living: - The **Federal Reserve** built the index and maintains it. - The **IMF** uses it in its [Global Financial Stability Report](https://www.imf.org/-/media/files/publications/gfsr/2025/april/english/ch2.pdf) to study how geopolitical shocks move stock prices, bond yields, and bank exposures, and in its [World Economic Outlook](https://www.imf.org/-/media/files/publications/weo/2026/april/english/ch1.pdf), where the country-level index tracks regional risk inside the global forecast. - The **European Central Bank and the European Systemic Risk Board** build it into financial-stability stress scenarios, bank-lending risk models, and growth-at-risk frameworks, in their [January 2026 fragmentation report](https://www.ecb.europa.eu/pub/pdf/other/ecb.report202601_financialstabilityrisks.en.pdf). The use cases are just as broad. The index has been used to forecast investment and growth, to design stress scenarios, to price equity and credit risk, to identify oil supply shocks, and to track capital flight from emerging markets. So the measure is valuable and trusted. The question we asked is whether the way it is built leaves room to do better. ## What we are matching There are two published versions, both from the economists who created this field, and we test against both. **The original Geopolitical Risk index** ( [Caldara and Iacoviello, 2018](https://www.federalreserve.gov/econres/ifdp/files/ifdp1222.pdf) ) counts geopolitical risk words across ten major newspapers, as a share of all articles published. It is the long-running benchmark, and it reaches back over a century. **The AI-GPR index** ( [Iacoviello and Tong, 2026](https://www.matteoiacoviello.com/research_files/AI_GPR_PAPER.pdf) ) is the modern successor. Instead of counting words, it asks a language model to read each article and score how much real geopolitical risk it carries, from zero to one. That removes most of the miscounting that word-matching suffers from, and it lets the authors add cuts the original never had: a score for each country, for each pair of countries, and for oil supply. The AI-GPR index is the harder target, and it is the one we hold ourselves to. It is the most accurate version, and it already publishes the country, pair, and oil breakdowns we want to reproduce. Matching those, from the data we already hold, is the bar we set. ## The solution: one number from the world's news The Nosible World database already does the hard part. Every event is tagged with its topic, its main country, the geopolitical entities it names (extracted by named-entity recognition), and how many separate publishers covered it. One real event is one record, no matter how many outlets repeat it. The database covers 2010 to today. This study uses the window from 2019 onward, where the published indices overlap, across 13.2 million de-duplicated events. That makes the index simple to build. Every event is tagged with one topic from an ontology over the events: the IPTC Media Topics ontology, a structured three-level hierarchy that classifies what each event is about. Its top level has buckets such as `conflict, war and peace`, `economy, business and finance`, and `health`, each splitting into finer levels below. We treat an event as geopolitical with an exact filter on three ontology fields carried by every event ( `iptc_level_1`, `iptc_level_2`, `iptc_level_3` ): ``` geopolitical(e) is TRUE when ANY of these hold: iptc_level_1 == "conflict, war and peace" iptc_level_2 == "international relations" iptc_level_3 in { "war crime", "genocide", "terrorism", "nuclear policy" } ``` The level-1 bucket is the dense spine (armed conflict, terrorism, coups, civil unrest); the level-2 and level-3 additions bring in cross-border politics and a few high-precision leaves that sit outside that bucket. We match the ontology codes exactly, never the words in the labels, so "weather warning" or "tug-of-war" never leak in. Write `breadth(e)` for the number of distinct publishers that covered an event `e`, which is the `total_netlocs` field carried on every event. The index, each day, is the share of total publisher attention spent on geopolitical events: ``` sum of breadth(e) over geopolitical events on day t NOSIBLE-GPR(t) = ----------------------------------------------------- sum of breadth(e) over all events on day t ``` Coverage breadth is the weight: a story carried by 200 outlets counts as one event weighted by 200, never as 200 separate events. For the charts we rescale each series to average 100 over 2020 to 2024, which makes them comparable and leaves every correlation unchanged. With Nosible World this is one group-by over the event table: filter to the conflict topics, sum `total_netlocs` for the numerator, and divide by the same sum over all events, per day. ## It matches the benchmark We put our index next to the published series, on the same scale. We show two versions of our signal. Both are the same share of publisher attention spent on geopolitical events, and differ only in what they divide by. The first divides by total attention that same day. The second divides by a trailing 12-month average of total attention. The second version exists because our news corpus grew significantly from 2019 to 2026, and it broadened across topics as it grew. So the share spent on any one theme drifts down over the years even when the world is no calmer, which quietly flattens the most recent period. The benchmark runs on a small, stable set of newspapers and has no such drift. Dividing by a trailing 12-month baseline removes ours, so 2026 is measured on the same footing as 2019. This detrended version is the one we carry through the rest of this post, for every country, every pair, and oil. ![Nosible geopolitical risk signal, raw daily share and 12-month detrended, against the published Federal Reserve and AI versions, 2019 to 2026](https://nosible.com/images/2026/06/nosible-gpr-vs-published.png) Nosible geopolitical risk signal, raw daily share and 12-month detrended, against the published Federal Reserve and AI versions, 2019 to 2026 Every event that should appear, appears: Soleimani in 2020, the invasion of Ukraine in 2022, Israel and Gaza in 2023, Iran and Israel in 2024, and the US-Iran war in 2026. Both versions track the AI-GPR index closely: the raw daily share at 0.90 on the levels and 0.79 on the stricter month-to-month changes, and the detrended version at 0.89 and 0.75. They are deliberately close, because the detrend changes the slope of the baseline, not which events are geopolitical. The chart shows where they part: the detrended line holds the 2026 surge up near the benchmark, while the raw share sags as the corpus swells. The two published versions agree with each other at about the same level, and we reach it on different data, with no keyword list. **What you get:** a live geopolitical risk series, updated automatically, with no analyst in the loop. ## Risk for every country You can also measure each country on its own, not just the world total. Each event already carries one main country: the country it is mainly about. But most geopolitical events involve more than one country, and attributing an event only to its main country misses the rest. An event whose main country is Lebanon is often just as much about Israel. An event whose main country is the United States may be about sanctions on Iran. So we use every country the event names, not just its main one. Each event already lists its geopolitical entities (countries, cities, regions) in the `ent_gpe` field, extracted by named-entity recognition. We map those entity strings to countries and attribute the event to all of them, which means handling their many forms: - Names and abbreviations (United States, U.S., USA). - Nationalities (American, Russian, Israeli). - Other languages and scripts (中国, Россия, ישראל). We match each entity whole-string and case-insensitive against a table of country names, official aliases, nationalities of four letters or more, and native-language names. A partial match never counts, so `Indiana` never resolves to India, and genuinely ambiguous strings such as a bare `Georgia` (the US state) are dropped. This is the difference between a weak score and a strong one. The United States is named in roughly a third of all geopolitical news, so attributing each event to its single main country alone misses most of where the US actually appears. Counting it everywhere it is named lifts its score from 0.45 to 0.83. Per country, the index is the same publisher-attention share, attributed to every country an event names: ``` attribution(e) = { the event's main country } + { countries resolved from its named entities } sum of breadth(e) over geopolitical events in month m where country c is in attribution(e) GPR(country c, month m) = ----------------------------------------------------------- B(m) ``` `B(m)` is one global denominator shared by every country: the trailing 12-month average of total monthly breadth across all events. We divide by this global figure, not by a country's own coverage, because in a crisis a country's own coverage spikes too and would cancel the signal; the trailing average also stops the corpus growing over time from masking real spikes. To build it: explode each event to its attribution set, group by country and month, and divide by `B(m)`. ![Per-country signal, same-day and 12-month detrended, against the published country indices for Russia, Israel, Ukraine, Iran, the USA, and India](https://nosible.com/images/2026/06/nosible-gpr-by-country.png) Per-country signal, same-day and 12-month detrended, against the published country indices for Russia, Israel, Ukraine, Iran, the USA, and India The results hold across the major actors: Iran 0.99, Israel 0.96, Ukraine 0.96, Russia 0.94. Seventy-six countries score above 0.60. The AI-GPR index needs an extra language-model pass over every article to do this. We get it from tags the data already carries. **What you get:** a risk monitor for any country, built the same way as the global one. ## Risk for every country pair Because each event already carries every country named in it, we also know which two appear together. That gives a risk score for any pair of countries. ![Bilateral signal, same-day and 12-month detrended, for major country pairs against the published bilateral series](https://nosible.com/images/2026/06/nosible-gpr-bilateral.png) Bilateral signal, same-day and 12-month detrended, for major country pairs against the published bilateral series From the same attribution sets, take every unordered country pair an event names: ``` sum of breadth(e) over geopolitical events in month m that name BOTH country a and country b GPR(pair a-b, month m) = ----------------------------------------------------------- B(m) ``` Same `B(m)` as the country index. To build it: self-join each event's attribution set into unordered pairs, then group by pair and month. The major conflict pairs come through clearly: Iran and the USA at 0.97, Russia and Ukraine at 0.95, India and Pakistan at 0.95. Country-pair risk is the newest part of the AI-GPR index. We reproduce it as a by-product of the country work. One pair does not match: China and the USA, at 0.28. The reason is precise, and it comes back to that one topic per event. A tariff story gets the topic "international trade," which sits in the economy branch of the ontology, not the conflict branch our filter reads, so the filter never sees it. We cannot just add the trade topic either, because it is mostly routine commerce and would drown the signal. This is a real limit, and it has a clean fix. We close it in the next section, and the score climbs from 0.28 to 0.75. **What you get:** a tension tracker for any two countries, useful for trade, supply chains, and exposure. ## Closing the trade-war gap The China-USA gap has a clean fix, and it does not touch the topic filter. The filter misses tariffs because each event carries only one topic, but every event also carries an embedding, a numeric summary of its meaning. We can read that directly. We wrote one phrase for trade coercion: a government imposing tariffs, duties, or export controls on another country, and the threats of retaliation that follow. We compute the cosine similarity between the phrase's embedding and every event's embedding to get a trade-coercion score, and let an event into the index if it is geopolitical *or* it clears that score. Nothing else changed: the same publisher-breadth weight, the same country attribution, the same denominator. ``` trade(e) = cosine similarity between event e's embedding and this trade-coercion anchor phrase: "Tariffs, trade wars and economic coercion between countries: a government imposing import tariffs, retaliatory duties, export controls or other trade restrictions on another country, and the diplomatic tensions and threats of retaliation these trigger" sum of breadth(e) over month m, country c in attribution(e), where geopolitical(e) OR trade(e) >= 0.40 GPR(country c, month m) = ------------------------------------------------------------------ B(m) ``` The only change from the country index is the `OR trade(e) >= 0.40`: an event now counts if it is geopolitical or it clears the trade-coercion score. Everything else is identical, which is why the conflict countries do not move. ![China and the China-USA pair, before and after folding in trade coercion. AI-GPR in amber, the Nosible baseline in grey, the Nosible version with tariffs in green](https://nosible.com/images/2026/06/nosible-tariff-recovery.png) China and the China-USA pair, before and after folding in trade coercion. AI-GPR in amber, the Nosible baseline in grey, the Nosible version with tariffs in green It closes the gap. China's country score rises from 0.53 to 0.78, and the China-USA pair rises from 0.28 to 0.75. It takes only 8,079 added events across the entire corpus, because one real trade-war event, weighted by how many outlets cover it, carries the spike. The conflict-driven countries do not move: Ukraine, Russia, Israel, and Iran stay where they were. The April 2025 tariff spike, missing before, now appears. **What you get:** the same method, pointed at a different gap, with no new model to train. ## Risk to oil supply Oil is the headline application in the AI-GPR paper, so we follow their method closely and compare to it directly. Oil needs care because the direction is not obvious. Broad geopolitical risk usually *lowers* the oil price, because fear cuts demand. Risk inside oil-producing regions *raises* it, because it threatens supply. A useful signal has to isolate the supply side. The paper does this in two steps. It keeps the geopolitical articles, filters them to the ones that mention oil, then asks a language model whether each one describes an oil supply disruption, and in which region. Our topic filter cannot do this on its own. Each event gets a single topic, so a Gulf war is tagged conflict, never energy. That one label can tell us a story is geopolitical, or that it is about oil, but never both at once. So we use the same second signal that closed the trade-war gap: the event embedding. We score each event against three oil-supply-risk phrases the same way, and keep the events that are both geopolitical and above the relevance floor. The embedding does the work of the paper's keyword filter and its disruption model in one step, with no extra model call. ``` relevance(e) = highest cosine similarity between event e's embedding and these three oil-supply-risk anchor phrases: "Armed conflict and war in or around major oil-producing regions" "Crude oil supply: OPEC production quotas, output levels and disruptions to oil production or exports" "Sanctions, embargoes and export restrictions on major oil-exporting countries such as Iran, Russia and Venezuela" sum of relevance(e) x breadth(e) over events in month m that are geopolitical AND have relevance(e) >= 0.30 Oil-GPR(m) = ---------------------------------------------------------------------- B(m) ``` To build it: score every event's stored embedding against the anchor phrases, keep the geopolitical events above the relevance floor, weight each by relevance times breadth, and divide by `B(m)`. The per-region and per-country versions attribute each surviving event to its producer regions or countries first, exactly like the country index. The paper's main oil chart plots its index against the real oil price. Here is ours, with the published academic version on the same axis. ![Nosible Oil-GPR, 12-month detrend in green and same-day in blue, with the published academic Oil-GPR in amber and the WTI oil price in grey. After Iacoviello and Tong, 2026, Figure 4](https://nosible.com/images/2026/06/nosible-oil-gpr-vs-wti.png) Nosible Oil-GPR, 12-month detrend in green and same-day in blue, with the published academic Oil-GPR in amber and the WTI oil price in grey. After Iacoviello and Tong, 2026, Figure 4 The index lines sit almost on top of each other. Our oil signal reproduces the published academic version closely under both denominators: the same-day share at 0.95, and the 12-month detrended version, the one we carry throughout, at 0.97. The paper also breaks the oil signal down by producer region. We reproduce that, and check each region against the academic version directly. ![Per-region oil-supply risk: the published academic version in amber, the Nosible version in green, by producer region, monthly and z-scored](https://nosible.com/images/2026/06/nosible-oil-gpr-by-region.png) Per-region oil-supply risk: the published academic version in amber, the Nosible version in green, by producer region, monthly and z-scored The major producers line up closely: the Middle East at 0.95, Venezuela at 0.96, the United States at 0.92, Russia at 0.91. Russia's signal jumps in 2022 with the invasion of Ukraine. Two regions stay weak: Africa at 0.44 and the North Sea at 0.04, where the published series is itself close to noise over this window. We are careful about one claim. We match the academic risk index, not the oil price itself. The paper goes further and shows that an oil-supply shock pushes the oil price up and output down. We leave demonstrating that from the Nosible signal to future work. **What you get:** an oil-supply risk signal that reproduces the academic benchmark, broken down by region. ## One method, many signals This is not one index. It is a template, and it runs on two engines. The first is the ontology over the events. The same recipe builds a signal for any subject the ontology classifies: - Economic policy uncertainty. - Stock-market volatility, against the VIX. - Climate concern and pandemic risk. The second is the meaning of the text. When a subject does not fit one topic, a phrase captures it instead, in every language at once. We have now done this twice, for oil supply and for trade coercion, with the same few lines of code. It extends to any traded asset: - Gold, natural gas, wheat, copper. - Freight and shipping. - Individual currencies. Each one is a use case: idea generation, risk monitoring, or a clean input for your own models. The number of signals you can build is effectively unlimited. We rebuilt one of the most widely used risk measures in finance and matched it across every cut its authors publish: the global index, the country breakdown, the country pairs, and oil supply. We did it from one multilingual database, with no language model reading each article. The same approach now points at every other index on the shelf. Nosible turns the world's news into a structured, multilingual, de-duplicated event database, and this post used one slice of it. If you want access to that database, or a signal like these built for your own models, [start a trial](https://nosible.com/start-trial). You can explore the live data at [nosible.world](https://nosible.world/). To inspect the same idea as an editable definition, [explore 30 worked semantic factor definitions built from NOSIBLE World](https://nosible.com/semantic-factors). ## References - Caldara, Dario and Matteo Iacoviello. [Measuring Geopolitical Risk](https://www.federalreserve.gov/econres/ifdp/files/ifdp1222.pdf). Board of Governors of the Federal Reserve System, International Finance Discussion Paper No. 1222 (2018); published version in the *American Economic Review* 112(4), 2022. Index and data: [matteoiacoviello.com/gpr](https://www.matteoiacoviello.com/gpr.htm). - Iacoviello, Matteo and Jonathan Tong (2026). [The AI-GPR Index: Measuring Geopolitical Risk using Artificial Intelligence](https://www.matteoiacoviello.com/research_files/AI_GPR_PAPER.pdf). Federal Reserve Board working paper. Overview and data: [matteoiacoviello.com/ai_gpr](https://www.matteoiacoviello.com/ai_gpr.html). - International Monetary Fund (2025). [Global Financial Stability Report, April 2025, Chapter 2: Geopolitical Risks and Their Implications for Asset Prices and Financial Stability](https://www.imf.org/-/media/files/publications/gfsr/2025/april/english/ch2.pdf). Measures geopolitical risk with the Caldara and Iacoviello (2022) indices, and estimates GPR betas, sovereign-yield and bank-exposure responses, and downside risk to stock returns. - International Monetary Fund (2026). [World Economic Outlook, April 2026, Chapter 1](https://www.imf.org/-/media/files/publications/weo/2026/april/english/ch1.pdf). Uses the global and country-specific Caldara and Iacoviello geopolitical risk indices to track regional risk and estimate the macroeconomic effects of geopolitical shocks. - European Central Bank and European Systemic Risk Board (2026). [Financial Stability Risks from Geoeconomic Fragmentation, January 2026](https://www.ecb.europa.eu/pub/pdf/other/ecb.report202601_financialstabilityrisks.en.pdf). Uses the Caldara and Iacoviello GPR index in its geopolitical-shock scenarios and VAR models. [All Research](https://nosible.com/blog) Related Research ![Thirty semantic factors rendered as daily time series on a dark grid, from geopolitical risk to firm-level political risk.](https://nosible.com/images/2026/08/semantic-factors-hero.png) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ### [The Future That Could Have Been: Turning Web Text into Semantic Stock Betas](https://nosible.com/blog/the-future-that-could-have-been-turning-web-text-into-semantic-stock-betas) 2026-08-05 11 min read ![Signed contrast number line separating systemic and idiosyncratic risk](https://nosible.com/images/2026/06/walkthrough_step9_number_line.png) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ### [Two Tricks for Turning Sentence Embeddings into Clean Features](https://nosible.com/blog/the-contrastive-geometry-of-risk) 2026-06-18 14 min read ![Daily NOSIBLE Trade Policy Uncertainty index compared with the published index](https://nosible.com/images/2026/06/nosible-tpu-daily.png) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ### [An Embedding-Based Approach to Trade and Economic Policy Uncertainty](https://nosible.com/blog/an-embedding-based-approach-to-trade-and-economic-policy-uncertainty) 2026-06-17 23 min read > Rebuild the geopolitical risk index from 13.2 million NOSIBLE World events and reproduce its global, country, country-pair and oil-risk signals. **URL:** https://nosible.com/blog/rebuilding-the-geopolitical-risk-index-from-nosible-world --- --- title: "Financial Sentiment" description: "Research on measuring, modelling, and validating financial-news sentiment." url: "https://nosible.com/blog/tag/financial-sentiment" --- /res Research # Financial Sentiment Research on measuring, modelling, and validating financial-news sentiment. [All](https://nosible.com/blog) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Financial Sentiment](https://nosible.com/blog/tag/financial-sentiment) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [Trading Signals](https://nosible.com/blog/tag/trading-signals) [Web Search](https://nosible.com/blog/tag/web-search) 3 articles ![Running sprinter illustration representing efficient financial-sentiment model training](https://nosible.com/blog/illustrations/the-sprinter.png) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [Financial Sentiment](https://nosible.com/blog/tag/financial-sentiment) ## [Matching GPT-5.1 at Financial Sentiment with Active Learning and Qwen3](https://nosible.com/blog/fast-enough-to-matter-productionizing-tiny-transformers-for-signal-extraction) Here's how we fine-tuned Qwen3 0.6B to beat FinBERT and match GPT-5.1 accuracy. Complete with open-source models, datasets, and training scripts. Spoiler alert: active learning is all you need. 2025-12-12 27 min read ![3D cube illustration representing LLM ensemble distillation](https://nosible.com/blog/illustrations/cube.png) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [Financial Sentiment](https://nosible.com/blog/tag/financial-sentiment) ## [A Pattern for Scaling the Value Proposition of LLMs: Ensemble and Distil 🚀](https://nosible.com/blog/ensemble-and-distil) We introduce the ensemble and distil data pattern and use it to fit an ordinary least squares linear regression that outperforms GPT-4 at financial news sentiment classification using sentence transformer embeddings as features. 2024-02-06 12 min read ![Abstract eye illustration representing financial-news sentiment analysis](https://nosible.com/blog/illustrations/eye.png) [Financial Sentiment](https://nosible.com/blog/tag/financial-sentiment) ## [News Sentiment Showdown: Who Checks Vibes Best?](https://nosible.com/blog/news-sentiment-showdown-who-checks-vibes-best) A comparison of sentiment classifications made by TextBlob, VADER, Flair, SigmaFSA, FinBERT, FinBERT-Tone, Text-Bison, Text-Unicorn, Gemini-Pro, GPT-3.5, GPT-4, and GPT-4-Turbo. We look at accuracy, time, and cost and include a dataset of 10,368 labelled news stories (with code) for our followers. 2024-01-28 12 min read [All Research](https://nosible.com/blog) > Research on measuring, modelling, and validating financial-news sentiment. **URL:** https://nosible.com/blog/tag/financial-sentiment --- --- title: "Matching GPT-5.1 at Financial Sentiment with Active Learning and Qwen3" description: "Fine-tune Qwen3 0.6B with active learning to match GPT-5.1 on financial sentiment, with open models, datasets and training code." url: "https://nosible.com/blog/fast-enough-to-matter-productionizing-tiny-transformers-for-signal-extraction" --- [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [Financial Sentiment](https://nosible.com/blog/tag/financial-sentiment) # Matching GPT-5.1 at Financial Sentiment with Active Learning and Qwen3 [Simon van Dyk](https://www.linkedin.com/in/simon-van-dyk/) 2025-12-12 27 min read Copy as Markdown ![Running sprinter illustration representing efficient financial-sentiment model training](https://nosible.com/blog/illustrations/the-sprinter.png) Every hedge fund we meet asks the same question: do you offer aspect-based financial sentiment? Why? Because sentiment moves markets. Especially when it's a forward-looking statement about a material event. They know it. We know it. It's how RavenPack grew into a 200+ person company. Today we're sharing how we built and open sourced a financial sentiment model that beats FinBERT [1](https://nosible.com/blog/fast-enough-to-matter-productionizing-tiny-transformers-for-signal-extraction#user-content-fn-1) and matches GPT-5.1's accuracy at a fraction of the cost. The results: **87.34% accuracy on real-world data** and **86.4% on Financial PhraseBank [2](https://nosible.com/blog/fast-enough-to-matter-productionizing-tiny-transformers-for-signal-extraction#user-content-fn-2)**, outperforming FinBERT on both datasets while running orders of magnitude faster and cheaper than frontier LLMs. Do we really need another sentiment model? Didn't FinBERT solve this in 2022? In a nutshell: no it's not solved, because while FinBERT performs well on the popular Financial PhraseBank dataset, it sucks on real-world data. The Financial PhraseBank dataset is contrived. It doesn't resemble what real-world data looks like, and models trained on it are overfit. LLMs, on the other hand, generalize well and are surprisingly good at labeling textual data, but they're intractably expensive to use at scale. Therefore, over the past few weeks, we have trained and productionized three text classifiers. This post explains how. One classifier predicts the financial sentiment of a text snippet. Another determines whether the text is a forward-looking statement or not. The third determines whether the text contains a prediction or not. Combining [these signals along with other dimensions in NOSIBLE's data](https://nosible.com/files/2025/12/tesla-signals-sample.xlsx) enables us to do powerful aspect based analysis to answer questions like: - Which retail companies show persistent negative sentiment in forward-looking statements about margin pressure, despite positive sentiment about revenue growth? - Show me the correlation between negative forward-looking statements about regulatory risk and subsequent stock volatility for pharma companies in Q3 2024. - How is sentiment diverging between forward-looking guidance versus actual results across semiconductor companies this quarter? We've [open sourced these models on HuggingFace](https://huggingface.co/NOSIBLE) along with the datasets they were trained on. We also share source code for your own projects. Wait, but you're a search engine, right? NOSIBLE is already a world-class search engine, but more importantly it's **incredibly fast**, which unlocks the ability to build [near real-time datafeeds](https://x.com/StuartReid1929/status/1920720746415849703). Our Search Feed product lets you turn any search about any topic into a point-in-time, backtest-friendly, deduplicated time series of data. [Here's an example for Tesla](https://nosible.com/files/2025/12/tesla-top-30.xlsx). We've leaned even harder in this direction. Our Search Feeds product is now better described as **web-scale surveillance**. To prove this isn't just a marketing term, we're extracting and including these signals in our Search Feeds for our customers, and **we're just getting started**. Let's get to it. ## Methodology Our approach combines three ingredients: real-world data, active learning for label refinement, and fine-tuning a small causal LLM for production inference. ### Why Real-World Data Generalizes Better A quality dataset is key to building a good ML model. The combination of compute, a return to MLPs, and a much larger labeled dataset contributed significantly to what made AlexNet so successful back in 2012. Using the vast amount of real-world data indexed by NOSIBLE, we sampled a Search Feed centered around financial news. The texts are varied and contain a ton of signal. Here's what they look like: > Some peers in the therapeutics segment have already reported their Q3 results. Gilead Sciences posted year-on-year revenue growth of 3%, beating expectations by 3.7%, and Biogen reported revenue up 2.8%, topping estimates by 8.2%. Following their reports, Gilead Sciences's stock traded up 1.2% and Biogen's stock was up 4.1%. There are fun ones too: > If you thought the wait for Kingdom Hearts III was ridiculous, you ain't see nothing yet. Now that we have an official release date for the game, Square Enix has already announced two collectors editions of the game. The first is your average run of the mill collectors edition, priced at $80 dollars. That's not the one you want. No, the big daddy collectors edition itself, the deluxe edition is the one you're going to want. This variety makes them difficult to label accurately. It's also what makes them excellent training data. The nuanced examples force models to learn robust patterns rather than memorize artifacts. Labeling 100,000 examples by hand is intractable. Instead, we used a combination of hand-labeling, LLM ensemble labeling, and active learning to produce a high-quality labeled dataset. ### Why Better Labels Are (Almost) All You Need ML practitioners can unintentionally obsess over model architecture and hyperparameters, but often all you need is better quality data. We took this approach: invest heavily in label quality, then starting with a simple baseline model, iteratively test larger and more complex models. The key insight is that correct labels on difficult examples matter more than thousands of mediocre labels. LLMs are excellent labelers, but they make mistakes on edge cases. Active learning helps you find and fix those mistakes systematically. Here's the four-step process: 1. **Look At Your Data**: Understand what makes classification difficult 2. **Label with LLMs**: Ensemble LLMs to label at scale 3. **Train Classifiers**: Build baselines to find signal 4. **Relabel Hard Texts**: Use active learning to improve label quality #### Step 1: Look At Your Data Don't vibe think. Go stare at the data. What makes it difficult? Sample 200 texts uniformly from your dataset and label them by hand: **negative**, **neutral**, or **positive**. This process isn't glamorous, but it's essential. You need to understand what makes classification difficult before you can build a good classifier. As you label, ask yourself: - What patterns distinguish **negative**, **neutral**, and **positive** sentiment? - Which examples are ambiguous or require domain knowledge? - Are any classes under-represented? Check your class distribution. Your sample doesn't need to match real-world distribution, it needs sufficient examples of each class for a model to learn from. If one class is under-represented (we needed more negative samples), sample more until you have enough examples to understand its patterns. This small investment pays dividends: you'll understand which edge cases will trip up your models later. The goal isn't just labels. It's understanding. What do easy-to-label samples look like? What makes the difficult ones difficult? This knowledge will guide your prompt engineering in the next step. #### Step 2: Label with LLMs Now that you understand your data, write a prompt YOU could follow to label it consistently. Find the smartest LLMs. Ensemble them together. Your prompt should be unambiguous. If you're labeling financial sentiment, define exactly what "negative," "neutral," and "positive" mean in your domain. Include examples of edge cases you discovered in Step 1. Make it a decision tree if that helps remove ambiguity. We used eight models: - [xAI: Grok 4 Fast](https://openrouter.ai/x-ai/grok-4-fast) - [xAI: Grok 4 Fast (reasoning enabled)](https://openrouter.ai/x-ai/grok-4-fast) - [Google: Gemini 2.5 Flash](https://openrouter.ai/google/gemini-2.5-flash) - [OpenAI: GPT-5 Nano](https://openrouter.ai/openai/gpt-5-nano) - [OpenAI: GPT-4.1 Mini](https://openrouter.ai/openai/gpt-4.1-mini) - [OpenAI: gpt-oss-120b](https://openrouter.ai/openai/gpt-oss-120b) - [Meta: Llama 4 Maverick](https://openrouter.ai/meta-llama/llama-4-maverick) - [Qwen: Qwen3 32B](https://openrouter.ai/qwen/qwen3-32b) Ensembling reduces individual model biases. To ensemble their labels, use majority vote. Start with your hand-labeled 200 samples to validate the prompt works before scaling to thousands. Here are the prompts we used for each classification task: Financial sentiment prompt ``` f""" # TASK DESCRIPTION Read through the following snippet of text carefully and classify the **financial sentiment** as either negative, neutral, or positive. You must also provide a short rationale for why you assigned the financial sentiment you did. For clarity here are the definitions negative, neutral, and positive sentiments: - **Negative**: The snippet describes an event or development that has had, is having, or is expected to have a material negative impact on the company's financial performance, share price, reputation, or outlook. - **Neutral**: The snippet is informational/descriptive and is not expected to have a material positive or negative impact on the company. - **Positive**: The snippet describes an event or development that has had, is having, or is expected to have a material positive impact on the company's financial performance, share price, reputation, or outlook. Materiality note: - “Material impact” includes likely effects on share price, revenue, costs, profitability, cash flow, guidance, regulatory exposure, reputation, risk exposure, or competitive position. # TASK GUIDELINES For the avoidance of doubt here is a decision tree that you can follow to arrive at the most appropriate sentiment classification for the snippet. Pay careful attention to the logic. Don't deviate. START │ ├── Step 1: Carefully read and understand the snippet. │ ├── Step 2: Check for sentiment indicators: │ ├── Is the snippet clearly NEGATIVE? │ (share price decline, losses, scandals, lawsuits, layoffs, product recalls, regulatory fines, │ leadership resignations, declining sales, market-share losses, reputational damage etc.) │ │ │ ├── YES → Classify as "negative" │ │ └── Provide a rationale by summarizing WHY the snippet is negative. │ │ │ └── NO → Continue below │ ├── Is the snippet clearly POSITIVE? │ (share price increases, strong earnings, favorable partnerships, successful product launches, awards, │ expansion plans, positive analyst coverage, reputational enhancement, etc.) │ │ │ ├── YES → Classify as "positive" │ │ └── Provide a rationale by summarizing WHY the snippet is positive. │ │ │ └── NO → Continue below │ └── If neither clearly positive nor negative → Classify as "neutral" (routine product announcements without performance implications, leadership appointments, scheduled reports, factual statements, general industry overviews, etc.) └── Provide a rationale by summarizing the WHY the snippet is neutral. If there is conflicting sentiment in the snippet pick the most dominant one, otherwise default to **neutral**. # RESPONSE FORMAT You must respond with ONLY a valid JSON object formatted as follows. DO NOT WRITE ANY PREAMBLE JUST RETURN JSON. {{ "rationale": "A one-sentence rationale for your classification", "financial_sentiment": "either negative, neutral, or positive" }} # SNIPPET TO LABEL Here is the snippet we would like you to assign a negative, neutral, or positive financial sentiment label to: {text} P.S. REMEMBER TO READ THE SNIPPET CAREFULLY AND FOLLOW THE GUIDELINES TO ARRIVE AT THE MOST APPROPRIATE FINANCIAL SENTIMENT CLASSIFICATION. WHEN IN DOUBT, YOU SHOULD DEFER TO A "neutral" CLASSIFICATION FOR THE SNIPPET. GOOD LUCK! """ ``` Forward-looking prompt ``` f""" You will be given a text snippet. Your task is to determine the **temporal orientation** of the main event or topic in the text, classifying it as either "forward" (forward-looking) or "not-forward" (backward-looking or neutral). # GUIDELINES For the avoidance of doubt here is a decision tree that you can follow to arrive at the most appropriate temporal orientation classification for the snippet. Pay careful attention to the logic. Don't deviate. START │ ├── Step 1: Carefully read and identify the MAIN event or topic in the text. │ (Ignore supporting details, commentary, or verb tenses of reporting) │ ├── Step 2: Determine the temporal orientation of this main event/topic: │ ├── Is the main event/topic FORWARD LOOKING? │ (Will the event occur in the future or is it planned/expected?) │ Examples: future launches, upcoming announcements, expansion plans, projections, │ forecasts, guidance, targets, goals, roadmaps │ Note: News about future plans (even if reported in past/present tense) = forward │ │ │ ├── YES → Classify as "forward" │ │ └── Provide a rationale explaining what future event the text focuses on. │ │ │ └── NO → Classify as "not-forward" │ └── This includes: │ • Past events (announcements made yesterday, completed mergers, reported earnings) │ • Current states (ongoing situations, present trading activity, existing conditions) │ • Timeless facts or general statements │ └── Provide a rationale explaining why the event is not forward-looking. │ END IMPORTANT: When uncertain about temporal orientation → Default to "not-forward" # DISAMBIGUATION RULES When the temporal orientation is unclear, apply these rules: 1. **Reporting Verb vs. Main Event Rule** - Ignore the tense of reporting verbs (said, announced, reported) - Focus on what is being reported about - Example: "CEO said yesterday the company will expand" → "forward" (expansion is future) 2. **Plans and Intentions Rule** - Any plans, intentions, targets, or forward guidance = "forward" (even if approved/decided in past) - Example: "Board approved new product launch" → "forward" (launch is future event) 3. **When in Doubt → Not-Forward** - If temporal orientation remains ambiguous → classify as "not-forward" # TEXT SNIPPET TO LABEL {text} You must respond with ONLY a valid JSON object formatted as follows. DO NOT WRITE ANY PREAMBLE JUST RETURN JSON. {{ "tense": "forward | not-forward", "rationale": "A one-sentence rationale for your classification" }} P.S. REMEMBER TO READ THE SNIPPET CAREFULLY AND FOLLOW THE GUIDELINES TO ARRIVE AT THE MOST APPROPRIATE CLASSIFICATION. WHEN IN DOUBT, YOU SHOULD DEFER TO A "not-forward" CLASSIFICATION FOR THE SNIPPET. GOOD LUCK! """ ``` Prediction prompt ``` f""" Read the following text snippet carefully and classify its **causal structure** as either **predictive** or **not-predictive**. You must also provide a short rationale for why you assigned the label you did. For clarity, here are the definitions: 1. Predictive: makes a concrete claim, forecast, prediction or estimate about a specific event. 2. Not Predictive: only reports or explains past or present facts. It does not contain ANY predictions, estimates, or forecasts. Plans, schedules, hopes, retrospectives or non-concreate predictions or estimates means it is **not-predictive**. For the avoidance of doubt, follow this decision tree exactly. Don’t deviate. START │ ├── Step 1: Carefully read and understand the snippet. │ ├── Step 2: Check for causal structure indicators: │ ├── Is the snippet clearly PREDICTIVE? │ (contains explicit forecasts or predictions about the future or │ numerical estimates about current or future events.) │ │ │ ├── YES → Classify as "predictive" │ │ └── Provide a rationale summarizing WHAT future outcome or effect is being │ │ forecast or expected. │ │ │ └── NO → Continue below │ └── If it is not clearly predictive → Classify as "not-predictive" └── Provide a rationale summarizing WHY it is 'not-predictive'. # DISAMBIGUATION RULES. 1. Predictive - If a snippet contains ANY predictive text, label it **"predictive"**. - Numerical estimates about CURRENT or FUTURE statistics are to be considered **predictive**. - Analyst estimates and ratings are **predictive** even if stated in the past tense. 2. Not predictive - Forward-looking plans or hopes without a claimed outcome remain **"not-predictive"**. - Schedules / announcements of events are not a forecast about an uncertain outcome or effect so **not-predictive**. - Retrospective narratives about past events should be classified as **not-predictive**. - Potential future actions without concrete predictions are **not-predictive**. - Future intent without a concrete prediction is **not-predictive**. You must respond with ONLY a valid JSON object formatted as follows. DO NOT WRITE ANY PREAMBLE JUST RETURN JSON. {{ "rationale": "A one-sentence rationale for your classification", "causal": "predictive | not-predictive" }} SNIPPET TO LABEL Here is the snippet we would like you to label: {text} P.S. REMEMBER TO READ THE SNIPPET CAREFULLY AND FOLLOW THE GUIDELINES TO ARRIVE AT THE MOST APPROPRIATE CAUSAL CLASSIFICATION. WHEN IN DOUBT, YOU SHOULD DEFER TO A "not-predictive" CLASSIFICATION FOR THE SNIPPET. GOOD LUCK! """ ``` Validate your prompt on your hand-labeled samples. Where do they disagree with the ensemble? Use those disagreements and the LLMs rationale to re-label examples and refine your prompt. We discovered earnings reports were frequently labeled **neutral** when they actually contained material outcomes, so we emphasized financial materiality over tone. Once validated, scale up to thousands of samples. Dataset transformations from 200 hand labeled samples to 10,000 LLM labeled samples ``` ┌────────┐ │ text │ shape (200, 1) │ str │ └────────┘ │ │ +1 human label ▼ ┌───────┬─────────────┐ │ text │ human_label │ shape (200, 2) │ str │ str │ └───────┴─────────────┘ │ │ +10,000 samples, +8 LLM labels, -1 Human label ▼ ┌───────┬─────────────┬─────┬───────────┐ │ text │ grok_4_fast │ ... │ qwen3_32b │ shape (10_000, 9) │ str │ str │ │ str │ └───────┴─────────────┴─────┴───────────┘ *10_200? Nope. │ We kept the samples for the human set separate │ +majority vote ▼ ┌───────┬─────────────┬─────┬───────────┬───────────────┐ │ text │ grok_4_fast │ ... │ qwen3_32b │ majority_vote │ shape (10_000, 10) │ str │ str │ str │ str │ str │ └───────┴─────────────┴─────┴───────────┴───────────────┘ ``` #### Step 3: Train Classifiers A baseline classifier is an accuracy reference, but can also tell you: - There is "signal": Verifies there are learnable features in the dataset. - Labels are good: Some features have a linear relationship with the labels we've assigned. - Benchmark for improvement: If a more complex model does poorly, you know the issue lies in something like the model architecture, training procedure, or a bug, not the dataset itself. Training classifiers in a nutshell: **1. Find a good representation:** Use text embeddings. Embedding models produce dense vectors that capture semantic meaning, including sentiment features. We ensembled six embedding models for more diverse representations. **2. Train baseline classifiers:** Start simple. Train a linear classifier (we used SGDClassifier with hinge loss) on your embeddings. If it performs well, you have signal. If it fails, investigate your features, labels, or task difficulty. **3. Check for overfitting:** Split your data (80/20 train/val). If training and validation accuracy are close, you're not overfit. Our baseline: 86.65% training accuracy, 85.93% validation accuracy. That's good signal with no overfitting. Text classification using embeddings ``` ┌───────────┐ ┌────────────┐ │ │ Embeddings │ │ Text │ Embedding │ representation │ Classifier │ Targets ─────────────▶ │ Model │ ────────────▶ │ Model │ ─────────▶ Label = "negative" ["Miss..."] │ │ [[0.42...]] │ │ [0,-1,...] └───────────┘ └────────────┘ Target mapping: {0: neutral, 1: positive, -1: negative} ``` The embedding models we used: - [Qwen3-Embedding-8B](https://openrouter.ai/qwen/qwen3-embedding-8b) - [Qwen3-Embedding-4B](https://openrouter.ai/qwen/qwen3-embedding-4b) - [Qwen3-Embedding-0.6B](https://openrouter.ai/qwen/qwen3-embedding-0.6b) - [OpenAI: Text Embedding 3 Large](https://openrouter.ai/openai/text-embedding-3-large) - [Google: Gemini Embedding 001](https://openrouter.ai/google/gemini-embedding-001) - [Mistral: Mistral Embed 2312](https://openrouter.ai/mistralai/mistral-embed-2312) Scaling to 100,000 samples added 1-2% accuracy giving us a final baseline: **86.65% training accuracy** and **85.93% validation accuracy** on the 20,000 validation samples. Code to train baseline classifier ``` import os import datetime as dt import polars as pl from sklearn.linear_model import SGDClassifier from sklearn.metrics import confusion_matrix, accuracy_score def embedding_model(data: pl.DataFrame, model: str, embeddings_path: str) -> tuple: """ Train a linear classifier on text embeddings for financial sentiment classification. Generates or loads embeddings, trains an SGDClassifier with hinge loss, and evaluates performance on train/validation splits. Returns predictions, accuracy metrics, and confusion matrices for both splits. :param data: DataFrame containing 'text' and 'majority_vote' columns with samples and labels :param model: Embedding model identifier for OpenRouter API :param embeddings_path: Path to cache/load embeddings (created if doesn't exist) :return: Returns a tuple of predictions, accuracy metrics, and confusion matrices for both splits. """ # Memoize the embeddings for subsequent training runs. if not os.path.exists(embeddings_path): # Generate the embeddings using OpenRouter and store as a numpy array file. embed_data(data=data, out_file=embeddings_path, model=model) # Load the embeddings, only as many samples as this training run. embeddings = np.load(embeddings_path) embeddings = embeddings[:data.shape[0]] # Fetch labels using majority vote and map to model targets. label_to_target={"positive": 1, "neutral": 0, "negative": -1} labels = list(label_to_target.keys()) targets = [label_to_target[row['majority_vote']] for row in data.to_dicts()] # Samples already shuffled, simple subscript split is sufficient. train_pct = 0.8 n_train = int(data.shape[0] * train_pct) x_train, y_train = embeddings[:n_train], targets[:n_train] x_val, y_val = embeddings[n_train:], targets[n_train:] # Get the training and validation text. train_text = data[:n_train].select("text").to_numpy().flatten().tolist() val_text = data[n_train:].select("text").to_numpy().flatten().tolist() classifier = SGDClassifier( loss='hinge', penalty='l2', alpha=0.00001, ) classifier.fit(x_train, y_train) # Make predictions. train_preds = classifier.predict(x_train) val_preds = classifier.predict(x_val) # Score the preds. train_acc = accuracy_score(y_true=y_train, y_pred=train_preds) val_acc = accuracy_score(y_true=y_val, y_pred=val_preds) confusion_train = confusion_matrix(y_true=y_train, y_pred=train_preds) confusion_train = { r: { c: int(confusion_train[i, j]) for j, c in enumerate(labels) } for i, r in enumerate(labels) } confusion_val = confusion_matrix(y_true=y_val, y_pred=val_preds) confusion_val = { r: { c: int(confusion_val[i, j]) for j, c in enumerate(labels) } for i, r in enumerate(labels) } print(f"{dt.datetime.utcnow()} - Train Accuracy {train_acc:.4f}") print(f"{dt.datetime.utcnow()} - Val Accuracy {val_acc:.4f}") print(f"{dt.datetime.utcnow()} - Confusion Matrix") print("Train", simplejson.dumps(confusion_train, indent=4)) print("Val", simplejson.dumps(confusion_val, indent=4)) return ( (y_train, train_preds, train_text), (y_val, val_preds, val_text), (train_acc, val_acc), (confusion_train, confusion_val) ) ``` A fun aside: including [some optional headers](https://openrouter.ai/docs/api/reference/overview#headers) in our embeddings requests to OpenRouter enabled them to measure our usage. Surprisingly, embedding a 100,000 sample dataset over 6 embedding models put us on the OpenRouter leaderboard for embeddings. Once we've operationalized this process into our Search Feeds product, we will sit in first place. In fact as of this writing we're in 2nd. OpenRouter leaderboard for Qwen3-Embedding-8B usage, November 2025 ![OpenRouter leaderboard showing NOSIBLE ranking 6th for Qwen3-Embedding-8B usage after the project](https://nosible.com/images/2025/12/openrouter-qwen3-embedding-8b-leaderboard.png) OpenRouter leaderboard showing NOSIBLE ranking 6th for Qwen3-Embedding-8B usage after the project #### Step 4: Relabel Hard Texts tldr; Use baselines to find difficult examples. Are they difficult? Or just wrong? Relabel the wrong ones. **What is Active Learning?** In traditional supervised learning, labels are fixed ground truth. But LLM generated labels aren't infallible. Active learning treats labels as mutable: find examples where your model struggles, investigate if the label is wrong, and fix it. Iterate until convergence. **Use baselines to find difficult examples:** Train your ensemble of linear models. Where do they all agree but the LLM ensemble label (majority vote) disagrees? These are your candidates for relabeling. When all your baseline models reach consensus, that's a strong signal the original label might be wrong. **Are they difficult? Or just wrong?** Not every disagreement means a bad label. Some examples are genuinely ambiguous. Consult a stronger LLM (we used OpenAI's GPT-5.1) as an oracle to make these decisions. If the oracle agrees with your baselines, relabel. If not, keep the original label. **Relabel iteratively:** Fix the labels, retrain your models, find new disagreements. The set of disagreements shrinks each iteration. Repeat until convergence. Iterate until no new disagreements appear. The basic outline of our full training algorithm with relabeling is as follows: 1. Label a set of 100k samples with the LLM labelers and compute majority vote. 2. Train multiple linear models on different embeddings of the same text to predict the majority vote of the LLM labelers. 3. Perform iterative relabeling: - Compare all the linear models' predictions to the majority vote label. - Identify disagreements where all linear models agree but the majority vote label does not. - Consult an oracle ([OpenAI: GPT-5.1](https://openrouter.ai/openai/gpt-5.1), a large LLM acting as our active-learning expert) to evaluate disagreements and relabel samples when appropriate. - Drop the worst performing linear model on the validation set from the ensemble. - Repeat until no additional samples require relabeling. 4. This is the final dataset used for training the classification models. The accuracy improvement was significant. ~3.5% Validation accuracy improvement as a result relabeling. | Split | Before | After | Improvement | | --- | --- | --- | --- | | Train | 86.65% | 90.03% | +3.38% | | Validation | 85.93% | 89.48% | +3.55% | Prompt used to consult the oracle ``` """ Consult the oracle to evaluate whether it agrees with the linear models prediction, or if we need to relabel this text sample. """ import textwrap labels = ["negative", "neutral", "positive"] labelling_prompt = textwrap.dedent(financial_sentiment_prompt(text)) label_options = " | ".join([f'"{label}"' for label in labels]) # The prediction variable is one of the labels above, where all linear models agreed. messages = [ { "role": "system", "content": [ { "type": "text", "text": textwrap.dedent( f""" # TASK DESCRIPTION Given the following LLM prompt, another model labelled this text as \"{prediction}\". Is this model correct? Provide the correct label and a rationale for your answer. # RESPONSE FORMAT You must respond using this exact format. DO NOT RESPOND IN ANY OTHER WAY: ```json {{ "correct": "true" | "false", "label": {label_options}, "reason": "A rationale for the correctness or incorrectness of your label." }} ``` P.S. Think fast there is no time to waste. """ ) } ] }, { "role": "user", "content": [{ "type": "text", "text": labelling_prompt }] } ] # Send these messages to OpenRouter using the OpenAI client. ``` ### Hijacking Qwen3 0.6B For Sentiment Analysis With 90% baseline accuracy and 100,000 clean labels, we had a production-quality dataset. Now for the final step: training a model fast and cheap enough to run at scale. We tested ModernBERT, DeBERTa, and several causal LLMs. The winner? Qwen3 0.6B, a tiny 600M parameter model that matched GPT-5.1's accuracy at orders of magnitude lower cost and latency. It's so small it's not even hosted on OpenRouter. You can run it on a phone. Why did Qwen3 0.6B work so well? Modern causal LLMs are pre-trained on massive text corpora, giving them strong language understanding out of the box. Fine-tuning them for classification is elegant: treat classification as next-token prediction. Given a prompt with the text snippet, the model predicts the label token ( **negative**, **neutral**, or **positive** ). No classification head, no special architecture, just next-token prediction. This means the representation changes too. Our baseline classifiers used embeddings, fine-tuning uses the tokenized text directly. The model learns to associate patterns in the raw token sequence with the correct label token. To ensure the model didn't overfit, we compared training loss with validation loss and monitored Financial PhraseBank accuracy during training. The implementation is straightforward, but three details are critical: **1. Label masking:** Only compute loss on the label token, not the prompt. This forces the model to learn classification, not memorize prompts. Set prompt tokens to -100 to tell PyTorch to ignore them: ``` labels = [-100] * len(prompt_tokens["input_ids"]) + answer_tokens["input_ids"] ``` **2. Left-padding:** Causal models need left-padding for batching. Right-padding makes the model attend to padding tokens when predicting, which breaks everything: ``` tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True) if tokenizer.pad_token is None: tokenizer.pad_token = tokenizer.eos_token tokenizer.padding_side = "left" ``` **3. EOS token handling:** During evaluation, exclude end-of-sequence tokens ( `<|im_end|>` for Qwen 0.6B) from accuracy calculations. Only measure whether the model predicted the correct label token (positive/negative/neutral), not conversational markup. We initially got suspiciously good results because we included EOS tokens, and fixing this gave us the true accuracy, which we verified after training on an out-of-sample dataset. To put this all together, here's the data flow for fine-tuning Qwen3 0.6B for financial sentiment classification: Fine-tuning data flow ``` ┌──────────────────────────────────────────────────────────┐ │ Training Example │ ├──────────────────────────────────────────────────────────┤ │ Text: "Tesla reported record deliveries..." │ │ Label: "positive" │ └──────────────────────────────────────────────────────────┘ │ ▼ ┌──────────────────────────────────────────────────────────┐ │ Build Messages │ ├──────────────────────────────────────────────────────────┤ │ System: "Classify the financial sentiment as positive, │ │ negative, or neutral." │ │ User: "Tesla reported record deliveries..." │ │ Assistant: "positive" (target label) │ └──────────────────────────────────────────────────────────┘ │ ▼ ┌──────────────────────────────────────────────────────────┐ │ Tokenize (apply chat template) │ ├──────────────────────────────────────────────────────────┤ │ Prompt tokens: [51234, 8273, ..., 19283] │ │ Answer tokens: [73421, 151645] (positive + EOS) │ └──────────────────────────────────────────────────────────┘ │ ▼ ┌──────────────────────────────────────────────────────────┐ │ Combine + Label Masking │ ├──────────────────────────────────────────────────────────┤ │ input_ids: [51234, 8273, ..., 19283, 73421, 151645] │ │ labels: [-100, -100, ..., -100, 73421, 151645] │ └──────────────────────────────────────────────────────────┘ │ ▼ ┌──────────────────────────────────────────────────────────┐ │ Qwen3 0.6B → Predict next tokens in sequence │ └──────────────────────────────────────────────────────────┘ │ ▼ ┌──────────────────────────────────────────────────────────┐ │ Loss computed only on non-masked tokens (73421, 151645) │ │ → Backprop → Update weights │ └──────────────────────────────────────────────────────────┘ ``` Here's our full training script ``` import os import datetime as dt import random import textwrap from itertools import chain import polars as pl import torch from datasets import Dataset, DatasetDict from sklearn.metrics import f1_score, accuracy_score from transformers import ( AutoModelForCausalLM, AutoTokenizer, DataCollatorForSeq2Seq, TrainingArguments, Trainer, AutoModelForSequenceClassification, DataCollatorWithPadding, ) from functools import partial def eval_finbert(): """ Evaluate the finbert model. :return: """ import numpy as np def compute_metrics(eval_pred): logits, labels = eval_pred predictions = np.argmax(logits, axis=-1) # Prefer Macro F1. f1 = f1_score(labels, predictions, average="macro") accuracy = accuracy_score(labels, predictions) return {"f1": f1, "accuracy": accuracy} finbert_model_id = "ProsusAI/finbert" finbert_label2id = { "positive": 0, "negative": 1, "neutral": 2, } finbert_id2label = {v: k for k, v in finbert_label2id.items()} finbert_num_labels = len(finbert_label2id) finbert_tokenizer = AutoTokenizer.from_pretrained(finbert_model_id) finbert_model = AutoModelForSequenceClassification.from_pretrained( finbert_model_id, num_labels=finbert_num_labels, label2id=finbert_label2id, id2label=finbert_id2label, ) def tokenize_finbert(batch): return finbert_tokenizer(batch["text"], truncation=True, max_length=512) # PhraseBank with FinBERT tokenizer. finbert_phrase_ds = fin_bank_ds.map(tokenize_finbert, batched=True) val_ds_for_finbert = val_ds.map(tokenize_finbert, batched=True) finbert_collator = DataCollatorWithPadding(tokenizer=finbert_tokenizer) finbert_eval_args = TrainingArguments( output_dir="finbert_baseline_eval", per_device_eval_batch_size=32, do_train=False, do_eval=True, logging_strategy="no", ) # Evaluate FinBERT on PhraseBank. finbert_trainer_phrase = Trainer( model=finbert_model, args=finbert_eval_args, data_collator=finbert_collator, eval_dataset=finbert_phrase_ds, compute_metrics=compute_metrics, ) print("\n---") print("Evaluating FinBERT on Financial PhraseBank...") finbert_phrase_metrics = finbert_trainer_phrase.evaluate() print("FinBERT metrics on PhraseBank:", finbert_phrase_metrics) print("---") # Evaluate FinBERT on 20% of data_labels. finbert_trainer_val = Trainer( model=finbert_model, args=finbert_eval_args, data_collator=finbert_collator, eval_dataset=val_ds_for_finbert, compute_metrics=compute_metrics, ) print("\n---") print("Evaluating FinBERT on 20% held-out data_labels...") finbert_val_metrics = finbert_trainer_val.evaluate() print("FinBERT metrics on 20% data_labels held-out set:", finbert_val_metrics) print("---\n") # Move the finbert model off the gpu. finbert_model.to(device="cpu") def compute_metrics_without_eos(eval_pred, eos_token_id): """ Compute the f1 score and the accuracy based on TOKEN classification. Given we are predicting a single token this is a good approximation of the actual scores. Make sure to remove eos token. :param eval_pred: Predictions. :param eos_token_id: The eos token id. :return: """ predictions, labels = eval_pred predictions = predictions[:, :-1] labels = labels[:, 1:] # Exclude Prompt (-100) AND EOS. valid_mask = (labels != -100) & (labels != eos_token_id) pred_flat = predictions[valid_mask] label_flat = labels[valid_mask] f1 = f1_score(label_flat, pred_flat, average="macro") accuracy = accuracy_score(label_flat, pred_flat) return {"f1": f1, "accuracy": accuracy} def preprocess_logits_for_metrics(logits, labels): """ Original logits are (Batch, Seq, Vocab). We only need (Batch, Seq) containing the indices of the max logit. """ if isinstance(logits, tuple): # Depending on the model and config, logits may contain extra tensors, # like past_key_values, but logits always come first logits = logits[0] return logits.argmax(dim=-1) if __name__ == "__main__": # ---------------------------------------------------- # STEP 1 - Load financial sentiment dataset # ---------------------------------------------------- task = "financial_sentiment" # Target TEXT strings for the Guard model. label_map = { "positive": "positive", "negative": "negative", "neutral": "neutral" } modeling_dir = os.path.join("/", "workspace", ".nosible") data_labels = pl.read_ipc(os.path.join(modeling_dir, f"{task}_100.0k_iter_18.ipc")) fin_bank = pl.read_ndjson(os.path.join(modeling_dir, "financial_phrase_bank.ndjson")) # Select text and text-labels. # Check if 'majority_vote' exists (it does in your IPC), otherwise rename or use 'labels'. if "majority_vote" in data_labels.columns: data = data_labels.select(["text", "majority_vote"]).rename({"majority_vote": "labels"}) else: data = data_labels.select(["text", "labels"]) fin_bank = fin_bank.select(["text", "labels"]) print(f"Data shape: {data.shape}") # ---------------------------------------------------- # STEP 2 - Create train/val splits # ---------------------------------------------------- # Slice to 100k only if we have that much data, otherwise take all. limit = min(100_000, len(data)) data = data[:limit] full_ds = Dataset.from_polars(df=data) fin_bank_ds = Dataset.from_polars(df=fin_bank) # Now this will succeed because we have at least 20 rows (16 train, 4 test) split_ds = full_ds.train_test_split(test_size=0.2, seed=42) train_ds = split_ds["train"] val_ds = split_ds["test"] ds = DatasetDict({ "train": train_ds, "val": val_ds, "phrasebank": fin_bank_ds, }) # ---------------------------------------------------- # STEP 3 - Fine-tune Qwen 3 (Guard Approach) # ---------------------------------------------------- # UPDATED: Use Qwen 3 (0.6B). model_id = "Qwen/Qwen3-0.6B" # Important: trust_remote_code=True is essential for Qwen3. # Make sure to set left padding. tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True) if tokenizer.pad_token is None: tokenizer.pad_token = tokenizer.eos_token tokenizer.padding_side = "left" # Load Model. model = AutoModelForCausalLM.from_pretrained( model_id, trust_remote_code=True, dtype=torch.bfloat16, device_map="auto" ) def tokenize(batch): """ Tokenize a batch. :param batch: Batch to tokenize. :return: """ input_ids_list = [] attention_mask_list = [] labels_list = [] for text, label in zip(batch['text'], batch['labels']): # 1. Build the prompt (User part) system = "Classify the financial sentiment as positive, negative, or neutral." msgs = [ {"role": "system", "content": system}, {"role": "user", "content": text}, ] prompt_str = tokenizer.apply_chat_template( msgs, tokenize=False, add_generation_prompt=True, enable_thinking=False, # By default, this is true. ) # 2. Build the answer (Assistant part) answer_str = f"{label}<|im_end|>" # Standard Qwen end token. # 3. Tokenize separately to know lengths prompt_tokens = tokenizer(prompt_str, add_special_tokens=False) answer_tokens = tokenizer(answer_str, add_special_tokens=False) assert len(answer_tokens["input_ids"]) == 2 # 4. Combine input_ids = prompt_tokens["input_ids"] + answer_tokens["input_ids"] attention_mask = prompt_tokens["attention_mask"] + answer_tokens["attention_mask"] # 5. CREATE LABELS WITH MASKING # -100 tells PyTorch to IGNORE these tokens during training. labels = [-100] * len(prompt_tokens["input_ids"]) + answer_tokens["input_ids"] # Truncate if necessary (simplified for brevity). if len(input_ids) > 2048: input_ids = input_ids[:2048] attention_mask = attention_mask[:2048] labels = labels[:2048] input_ids_list.append(input_ids) attention_mask_list.append(attention_mask) labels_list.append(labels) return { "input_ids": input_ids_list, "attention_mask": attention_mask_list, "labels": labels_list } # Tokenize the datasets. t_ds = ds.map(tokenize, batched=True, remove_columns=ds["train"].column_names, num_proc=8) # Set the collator for padding. data_collator = DataCollatorForSeq2Seq(tokenizer=tokenizer, padding=True) timestamp = dt.datetime.now().strftime('%Y%m%d_%H%M%S') task_output = f"{task}_qwen3_0.6B_{timestamp}" training_args = TrainingArguments( output_dir=task_output, # Qwen3 0.6B is very small, we might be able to increase batch size slightly. gradient_accumulation_steps=4, per_device_train_batch_size=16, per_device_eval_batch_size=16, num_train_epochs=5, learning_rate=2e-5, lr_scheduler_type="cosine", warmup_ratio=0.03, max_grad_norm=1.0, weight_decay=0.1, # Improvements for instruction fine-tuning. neftune_noise_alpha=5, # Optimizations. bf16=True, optim="adamw_torch_fused", logging_strategy="steps", logging_steps=10, logging_dir=task_output, logging_first_step=True, group_by_length=True, torch_compile=False, # Eval params. eval_strategy="steps", eval_steps=500, save_total_limit=3, load_best_model_at_end=True, metric_for_best_model="eval_val_loss", report_to="none", save_strategy="steps", save_steps=500, ) n_eval_train = int(0.05 * len(t_ds["train"])) chat_end_token_id = tokenizer.convert_tokens_to_ids("<|im_end|>") trainer = Trainer( model=model, args=training_args, data_collator=data_collator, train_dataset=t_ds["train"], eval_dataset={ "train": t_ds["train"].select(list(range(n_eval_train))), "val": t_ds["val"], "phrasebank": t_ds["phrasebank"], }, compute_metrics=partial(compute_metrics_without_eos, eos_token_id=chat_end_token_id), preprocess_logits_for_metrics=preprocess_logits_for_metrics, ) trainer.train() ``` ## Results How did our fine-tuned Qwen3 0.6B compare to FinBERT and frontier LLMs? ### Accuracy on Financial PhraseBank (Fake News) We compare the accuracy of our fine-tuned Qwen3 0.6B model against FinBERT and a variety of LLMs prompted for zero-shot classification on the Financial PhraseBank dataset by Malo, P. et al. (2014). This dataset contains 4,840 financial news snippets labeled for sentiment by human experts. ![Financial sentiment accuracy comparison on PhraseBank dataset](https://nosible.com/images/2025/12/financial_sentiment_accuracy_on_phrasebank_dataset.png) Financial sentiment accuracy comparison on PhraseBank dataset ### Accuracy on Real-World Data (Real News) Crucially, our model performs well on real world data too, not just benchmark datasets. Here are the accuracy results on the out of sample validation set from our 100,000 dataset. ![Financial sentiment accuracy comparison on NOSIBLE financial sentiment dataset](https://nosible.com/images/2025/12/financial_sentiment_accuracy_on_nosible_fin_sent_dataset.png) Financial sentiment accuracy comparison on NOSIBLE financial sentiment dataset ### Cost vs Accuracy tradeoff Finally, we compared the accuracy vs cost tradeoff of our model. This is important because while large LLMs can achieve high accuracy, their inference costs can be prohibitive at scale. Our fine-tuned model strikes a good balance between accuracy and cost. Although GPT-5.1 topped us by a margin, it's orders of magnitude more expensive, so in practical terms it's not feasible at our (NOSIBLE's) scale. ![Accuracy vs Cost plot: Financial sentiment accuracy on Financial PhraseBank vs Cost](https://nosible.com/images/2025/12/financial_sentiment_accuracy_vs_cost_on_phrasebank_dataset.png) Accuracy vs Cost plot: Financial sentiment accuracy on Financial PhraseBank vs Cost![Accuracy vs Cost plot: Financial sentiment accuracy on NOSIBLE financial sentiment dataset vs Cost](https://nosible.com/images/2025/12/financial_sentiment_accuracy_vs_cost_on_nosible_dataset.png) Accuracy vs Cost plot: Financial sentiment accuracy on NOSIBLE financial sentiment dataset vs Cost **Cost calculation methodology:** **For the LLMs:** We summed input and output token costs per million in a 10:1 ratio, which reflects the typical input/output length relationship when prompted with our labeling prompt. **For the fine-tuned model:** Since Qwen3 0.6B isn't hosted on OpenRouter, we estimated cost at a conservative 100:1 ratio given the minimal output—just the label token and EOS token. ## Conclusion Circling back to the question "do we really need another sentiment model?": #### FinBERT didn't solve financial sentiment in 2022 FinBERT, the established benchmark model, performs poorly on our real-world data despite strong performance on Financial PhraseBank. This suggests overfitting to cleaner, more structured benchmark datasets rather than the messy, varied text found in actual web data. #### LLMs are good at sentiment but don't scale Large LLMs like GPT-5.1 excel at financial sentiment classification but are prohibitively expensive at production scale. We show they are orders of magnitude more costly than our fine-tuned Qwen3 0.6B model at similar accuracy levels. #### Active learning makes auto-labeling possible at scale Active learning with our relabeling algorithm is key to creating a large, real-world and high-quality dataset from which models can learn effectively. One final insight from the project: causal models aren't just for text generation. When you understand how to leverage [logprobs](https://cookbook.openai.com/examples/using_logprobs), we show they can be powerful, production ready classification models too. ### Try it out! As a reminder, we've open sourced everything you need to build this yourself: - Financial Sentiment: - [NOSIBLE Financial Sentiment](https://huggingface.co/datasets/NOSIBLE/financial-sentiment) dataset - [NOSIBLE Financial Sentiment v1.1 Base](https://huggingface.co/NOSIBLE/financial-sentiment-v1.1-base) model - Forward Looking: - [NOSIBLE Forward-Looking](https://huggingface.co/datasets/NOSIBLE/forward-looking) dataset - [NOSIBLE Forward-Looking v1.2 Base](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) model - Prediction: - [NOSIBLE Prediction](https://huggingface.co/datasets/NOSIBLE/prediction) dataset - [NOSIBLE Prediction v1.1 Base](https://huggingface.co/NOSIBLE/prediction-v1.1-base) model **Quickstart on HuggingFace** 1. Visit our [model page](https://huggingface.co/NOSIBLE/financial-sentiment-v1.1-base) 2. Click **Deploy** -> **HF Inference Endpoints** 3. Configure the endpoint: - For Hardware choose a beefy GPU like a L40S - For Inference Engine choose SGLang and set Max Prefill Tokens: `65536` - For Container add container args: `--dtype float16 --cuda-graph-max-bs 128 --disable-radix-cache` - For Advanced Settings choose Download Pattern: **Download Everything** 4. Copy your endpoint URL and bring your HuggingFace API KEY. 5. Use the code below to classify some text. 6. Profit 🐒 Client code to interact with your HF Inference Endpoint. ``` import math from openai import OpenAI # Initialize the client pointing to your vLLM server client = OpenAI( base_url="YOUR_ENDPOINT_URL_HERE/v1", api_key="YOUR_API_KEY_HERE" ) model_id = "NOSIBLE/financial-sentiment-v1.1-base" # Input text to classify text = "The company reported a record profit margin of 15% this quarter." # Define the classification labels labels = ["positive", "negative", "neutral"] # Prepare the conversation messages = [ {"role": "system", "content": "Classify the financial sentiment as positive, neutral, or negative."}, {"role": "user", "content": text}, ] # Make the API call chat_completion = client.chat.completions.create( model=model_id, messages=messages, temperature=0, max_tokens=1, stream=False, logprobs=True, # Enable log probabilities to calculate confidence top_logprobs=len(labels), # Ensure we capture logprobs for our choices extra_body={ "chat_template_kwargs": {"enable_thinking": False}, # Must be set to false. "regex": "(positive|neutral|negative)", }, ) # Extract the response content response_label = chat_completion.choices[0].message.content # Extract the logprobs for the generated token to calculate confidence first_token_logprobs = chat_completion.choices[0].logprobs.content[0].top_logprobs print(f"--- Classification Results ---") print(f"Input: {text}") print(f"Predicted Label: {response_label}\n") print("--- Label Confidence ---") for lp in first_token_logprobs: # Convert log probability to percentage probability = math.exp(lp.logprob) print(f"Token: '{lp.token}' | Probability: {probability:.2%}") ``` ## Acknowledgments The team involved in the project includes: - [**Matthew Dicks**](https://www.linkedin.com/in/matthewdicks98/) - [**Simon van Dyk**](https://www.linkedin.com/in/simon-van-dyk/) - [**Gareth Warburton**](https://www.linkedin.com/in/garethwarburton/) - [**Stuart Reid**](https://www.linkedin.com/in/stuartgordonreid/) ## Citations **Datasets:** - [**NOSIBLE Financial Sentiment:**](https://huggingface.co/datasets/NOSIBLE/financial-sentiment) Nosible Inc. - [**Financial PhraseBank:**](https://huggingface.co/datasets/takala/financial_phrasebank) Malo, P. et al. (2014). **Research:** - **Qwen3:** *Qwen3 Technical Report*. [arXiv:2505.09388](https://arxiv.org/abs/2505.09388) - **Qwen3 Guard:** *Qwen3Guard Technical Report*. [arXiv:2510.14276v1](https://arxiv.org/html/2510.14276v1) **NOSIBLE Search Feed Examples:** - [Tesla Feed sample top-30](https://nosible.com/files/2025/12/tesla-top-30.xlsx) - [Tesla Feed signals sample](https://nosible.com/files/2025/12/tesla-signals-sample.xlsx) **Footnotes:** ## Footnotes 1. **FinBERT:** Araci, D. (2019). *FinBERT: Financial Sentiment Analysis with Pre-trained Language Models*. [arXiv:1908.10063](https://arxiv.org/abs/1908.10063) [↩](https://nosible.com/blog/fast-enough-to-matter-productionizing-tiny-transformers-for-signal-extraction#user-content-fnref-1) 2. **Financial PhraseBank:** [Malo, P. et al. (2014)](https://www.researchgate.net/publication/251231364_FinancialPhraseBank-v10) [↩](https://nosible.com/blog/fast-enough-to-matter-productionizing-tiny-transformers-for-signal-extraction#user-content-fnref-2) [All Research](https://nosible.com/blog) Related Research ![3D cube illustration representing LLM ensemble distillation](https://nosible.com/blog/illustrations/cube.png) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [Financial Sentiment](https://nosible.com/blog/tag/financial-sentiment) ### [A Pattern for Scaling the Value Proposition of LLMs: Ensemble and Distil 🚀](https://nosible.com/blog/ensemble-and-distil) 2024-02-06 12 min read ![High-resolution chart of selected 35B-A3B serving records across a 350-plus experiment campaign, from Qwen 3.5 FP8 on H200 at $0.218400 per million completion tokens in Experiment 006 to Qwen 3.6 source NVFP4 at $0.125959 in Experiment 332](https://nosible.com/images/2026/08/qwen-35b-a3b-serving-campaign-hero.svg) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) ### [What 350+ Experiments Taught Us About Serving Qwen 3.6 35B-A3B](https://nosible.com/blog/what-350-experiments-taught-us-about-serving-qwen-3-6) 2026-08-17 15 min read ![Thirty semantic factors rendered as daily time series on a dark grid, from geopolitical risk to firm-level political risk.](https://nosible.com/images/2026/08/semantic-factors-hero.png) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ### [The Future That Could Have Been: Turning Web Text into Semantic Stock Betas](https://nosible.com/blog/the-future-that-could-have-been-turning-web-text-into-semantic-stock-betas) 2026-08-05 11 min read > Fine-tune Qwen3 0.6B with active learning to match GPT-5.1 on financial sentiment, with open models, datasets and training code. **URL:** https://nosible.com/blog/fast-enough-to-matter-productionizing-tiny-transformers-for-signal-extraction --- --- title: "Can Faceted Search at Web-Scale Self Organize?" description: "See how adaptive named-entity tagging lets a web-scale search index organize itself into useful facets without relying on a fixed taxonomy." url: "https://nosible.com/blog/can-faceted-search-at-web-scale-self-organize" --- [Web Search](https://nosible.com/blog/tag/web-search) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) # Can Faceted Search at Web-Scale Self Organize? [Stuart Reid](https://www.linkedin.com/in/stuartgordonreid/) 2025-10-16 7 min read Copy as Markdown ![Abstract vortex illustration representing self-organizing web-scale search facets](https://nosible.com/blog/illustrations/vortex.png) ## Introduction Few things pack more punch than [faceted search](https://link.springer.com/book/10.1007/978-3-031-02262-3). Faceted search involves tagging documents in your index with various "facets". Facets could be content categories, geographies, industries, companies, individuals, etc. Then, when a query arrives, you map the query to the most relevant facets and search within them for the most relevant documents. If your classifier is good and your index supports pre-retrieval, you can unlock higher precision and lower latency. This is a common idea in Ecommerce search engines, and it works quite well... but one glaring issue exists: *what happens when facets are unbounded?* Take companies for example. Every day new companies are created by means of formations, mergers, demergers, joint ventures, spin-offs, carve-outs, etc. In my opinion, an intelligent search engine – especially one built for finance – should be capable of automatically discovering and understanding these events and updating its facets to match reality. It should self-organize. We've been iterating on this problem for a while and I'm excited to share our solution today. But before we get to that, I need to quickly recap [how our search engine is architected because](https://nosible.com/blog/the-road-to-cybernaut-1), frankly, it is different to most other search engines. ## We Are Many We built a deep learning model that learns how to segment web-scale corpora into near uniformly sized collections that are semantically and lexically coherent. Our last run learned 250,000 collections from 250,000,000 documents. Each of these collections are turned into their own miniature search engine. They have their own [vector index](https://en.wikipedia.org/wiki/Vector_database), [lexical index](https://en.wikipedia.org/wiki/Okapi_BM25), [rerankers](https://www.mongodb.com/resources/basics/artificial-intelligence/reranking-models?msockid=2fb6df0f720163f20ff5caf2730162e0), classifiers, and more. When a query arrives, we find the best collections using statistical methods and route the query to them. The results from each collection are fused and the best results are returned back to the user. This is unconventional but has massive benefits... like, for exampe, getting big model performance from tiny models. Every new document we index is written to the collection it belongs to and every so often those writes are flushed to disk. Larger or most important collections are flushed more frequently and are also more likely to be cached in RAM. ## Distilling Agents Coming back to named entity recognition, here's a cold hard truth: ***AI agents crush***. They can easily reason about named entities like a person does and using anything else today feels kind of like burning a CD - prehistoric. Unfortunately, AI agents are too slow and too expensive. At the time of writing we are indexing 15,000,000 new web pages every day. That works out to 91 million paragraphs a day, 180 million sentences, and 4.25 billion tokens. Passing all of that to an AI Agent every day is not tractable. Fortunately, we have found an elegant solution that gets us agentic-level tagging without the exorbitant costs and untenable latencies. Our solution is as follows. - Each collection contains two taggers: - An ultra-fast and precise tagger created for that collection using the outputs of the Resolver. This tagger is not neural, it uses Aho-Corasick. It can process data going into a collection in real-time. - A slow, universal tagger trained to identify named entities in general. This is a neural model, and it sees X% of the data added to each collection. X can be scaled to match compute capacity. - The slow, universal tagger suggests entities for each collection. Those suggestions are accumulated in a small database and, once a certain threshold is met, the suggestion is sent to the Resolver Agent. - The Resolver Agent uses a bunch of tools (Search, Wikipedia, LinkedIn, Market APIs, etc.) to work out who this entity is. This step is obviously slow and expensive, but it's also extremely cache friendly. - If the Resolver Agent accepts the new entity, a new set of patterns is created and sent to the collection(s) that suggested the entity. - Finally, when the collection is flushed it will check for new patterns. If new patterns are found it will tag all untagged chunks and update the patterns associated with that collection. This happens in seconds. The pattern is illustrated below: ![Self-organizing search facets arranged as a network around a central query](https://nosible.com/images/2025/10/image-1.png) Self-organizing search facets arranged as a network around a central query ## Why Does It Work? All of the intelligence exhibited by the Resolver Agent exists in the tokens. After all, the Resolve Agent is "merely" doing next token prediction. Logic dictates that it should therefore be possible to distil a great deal of that intelligence into a simple list of strings and regex patterns. This is not tractable in the general case of course, but it certainly works in specific cases like, for example, JP Morgan. Here is the data extracted from the web by the Resolver Agent: ``` { "name_orig": "JPMorgan Chase", "name_kgid": "JPMorgan Chase", "name_wiki": "JPMorgan Chase", "name_figi": "JPMORGAN CHASE & CO", "name_eodhd": "JPMorgan Chase & Co", "name_lei": "JPMORGAN CHASE & CO.", "ticker_wiki": "JPM", "ticker_figi": "JPM", "ticker_eodhd": "JPM.US", "ticker_serp": "JPM", "kgid": "kg:/m/01hlwv", "wiki_qid": "Q192314", "cik_code": "0000019617", "isin_code": "US46625H1005", "cusip_code": "16161A108", "lei_code": "8I5DZWZKVSZI1NUHU748", "figi_code": "BBG000DMBXR2", "open_corp_code": "us_de/691011", "ein_code": "13-2624428", "exchange": "New York Stock Exchange", "mic_code": "XNYS", "exch_code": "UN", "ccy_name": "US Dollar", "ccy_code": "USD", "ccy_symbol": "$", "wiki_page": "JPMorgan_Chase", "website": "http://www.jpmorganchase.com/", "linkedin": "https://www.linkedin.com/company/jpmorganchase", "is_company": true, "is_public": true, "is_private": false, "is_delisted": false, "continent": "North America", "region": "North America", "country": "United States", "country_iso": "US", "city": "New York City", "address": "383 Madison Avenue, New York, NY, United States, 10179", "phone_num": "(212) 270-6000", "sector": "Financials", "industry_group": "Banks", "industry": "Banks", "sub_industry": "Diversified Banks", "historical_names": [ "Chase Manhattan International Limited", "J.P. Morgan and Company", "Morgan Guaranty Trust Company", "Drexel, Morgan & Co.", "Chemical Banking Corporation", "Chemical Bank of New York", "The New York Chemical Manufacturing Company", "Bank One Corporation", "Banc One Corporation", "First Bancgroup of Ohio", "The Bank of the Manhattan Company" ], "brand_names": [ "JPMorganChase", "Chase", "Chase UK", "Chase Student Loans", "IndexGPT", "Quorum", "JPM Coin", "J.P. Morgan Workplace solutions" ], "subsidiaries": [ "Chase Bank", "JPMorgan Securities, LLC", "JPMorgan Europe, Ltd.", "Hambrecht & Quist", "Robert Fleming & Co.", "Texas Commerce Bank", "First Chicago NBD", "First Chicago Bank", "Banc One", "City National Bank of Columbus", "Purdue National Corporation", "Bear Stearns", "Washington Mutual", "First Republic Bank", "Collegiate Funding Services", "ClimateCare", "J.P. Morgan Cazenove", "Global Shares", "Renovite Technologies", "Viva Wallet", "Chase Manhattan Bank", "J.P. Morgan & Co.", "Bank One", "Chemical Bank", "Manufacturers Hanover", "National Bank of Detroit", "Providian Financial", "Great Western Bank", "Chase National Bank", "Corn Exchange Bank", "Guaranty Trust Company of New York", "JPMorgan Ventures Energy Corporation" ], "short_wiki": "TRUNCATED", "summary_raw": "TRUNCATED", "summary_wiki": "TRUNCATED", "summary_eodhd": "TRUNCATED", "markdown_wiki": "TRUNCATED", "debug_verbose": true } ``` And here are the unigrams, bigrams, and trigrams distilled from the information the Resolution Agent extracted. We would use these for tagging: ``` [ "jpm", "jpmorgan", "jpmorganchase", "j.p. morgan", "chase", "jpmorgan.com", "chase.com", "jpmcoin", "indexgpt", "dimon", "jpmorgan chase", "j.p. morgan", "jp morgan", "j p morgan", "chase bank", "nyse jpm", "jpmorgan securities", "jpmorgan cazenove", "jpmorgan europe", "jpmorgan ventures", "jpmorgan wealth", "jpmorgan private", "jpmorgan asset", "jpmorgan investment", "jpmorgan workplace", "chase uk", "chase sapphire", "chase freedom", "chase ink", "chase student", "chase ultimate", "bear stearns", "washington mutual", "first republic", "bank one", "chemical bank", "manufacturers hanover", "jpmorgan chase & co", "jpmorgan chase and co", "jp morgan chase", "j p morgan chase", "jpmorgan chase bank", "jpmorgan securities llc", "jpmorgan private bank", "jpmorgan asset management", "jpmorgan wealth management", "jpmorgan investment bank", "jpmorgan workplace solutions", "chase sapphire reserve", "chase sapphire preferred", "chase ultimate rewards", "chase credit cards", "chase retail banking", "house of morgan", "jamie dimon ceo", "jpmorgan corporate challenge", "chase center", "chase bank", "jpmorgan securities llc", "jpmorgan europe ltd", "hambrecht & quist", "robert fleming & co", "texas commerce bank", "first chicago nbd", "first chicago bank", "banc one", "city national bank of columbus", "purdue national corporation", "bear stearns", "washington mutual", "first republic bank", "collegiate funding services", "climatecare", "j.p. morgan cazenove", "global shares", "renovite technologies", "viva wallet", "chase manhattan bank", "j.p. morgan & co", "bank one", "chemical bank", "manufacturers hanover", "national bank of detroit", "providian financial", "great western bank", "chase national bank", "corn exchange bank", "guaranty trust company of new york", "jpmorgan ventures energy corporation" ] ``` ## Many Benefits Here are some benefits of this approach: 1. Not all entities exist in every collection. And the more coherent your collections, the truer this becomes. In fact, the distribution of entities is extremely sparse in the average case. So, learning collection-specific models saves a tonne of compute cost and minimizes spurious matches. 2. The approach is adaptive. We do not start with a hard coded list of entities. We allow each collection to organically discover and suggest new entities for itself thereby allowing it to adapt to change - formations, mergers, demergers, joint ventures, spin-offs, carve-outs, etc. 3. We unlock the benefits of agentic search without the prohibitive price tag. Originally, we planned to train our own NER models but, like I said, that kind of feels like burning a CD. Agentic search is here and using anything other than that for this task feels backwards. LMK if you'd like to see the Resolver Agent in action. Once we have decoupled it from our internal LLM generation tools we will open source it. [All Research](https://nosible.com/blog) Related Research ![NOSIBLE World knowledge graph showing entity connections over a decade](https://nosible.com/images/2026/07/kg-hero-decade.png) [Web Search](https://nosible.com/blog/tag/web-search) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ### [Point-in-Time Knowledge Graphs over Named Entities with NOSIBLE World](https://nosible.com/blog/point-in-time-knowledge-graphs-over-named-entities) 2026-07-16 8 min read ![Cybernaut-1 illustration representing agentic search with Monte Carlo Tree Search](https://nosible.com/blog/illustrations/cyber.png) [Web Search](https://nosible.com/blog/tag/web-search) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) ### [Introducing Cybernaut-1: Agentic Search using MCTS](https://nosible.com/blog/introducing-cybernaut-1-agentic-search-with-mcts) 2025-08-26 2 min read ![Railway-track illustration representing the road to Cybernaut-1](https://nosible.com/blog/illustrations/track.png) [Web Search](https://nosible.com/blog/tag/web-search) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) ### [The Road to Cybernaut-1: Rebuilding Search for AI](https://nosible.com/blog/the-road-to-cybernaut-1) 2025-08-20 17 min read > See how adaptive named-entity tagging lets a web-scale search index organize itself into useful facets without relying on a fixed taxonomy. **URL:** https://nosible.com/blog/can-faceted-search-at-web-scale-self-organize --- --- title: "Introducing Cybernaut-1: Agentic Search using MCTS" description: "Meet Cybernaut-1, NOSIBLE’s agentic search system combining hybrid retrieval with LLM-guided Monte Carlo Tree Search." url: "https://nosible.com/blog/introducing-cybernaut-1-agentic-search-with-mcts" --- [Web Search](https://nosible.com/blog/tag/web-search) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) # Introducing Cybernaut-1: Agentic Search using MCTS [Stuart Reid](https://www.linkedin.com/in/stuartgordonreid/) 2025-08-26 2 min read Copy as Markdown ![Cybernaut-1 illustration representing agentic search with Monte Carlo Tree Search](https://nosible.com/blog/illustrations/cyber.png) Today I am proud to announce the release of Cybernaut-1. Cybernaut-1 combines our powerful hybrid-3 search algorithm with LLM-guided [Monte Carlo Tree Search](https://en.wikipedia.org/wiki/Monte_Carlo_tree_search) to deliver world class search results on difficult queries. Cybernaut-1 is available via our [V2 Search API](https://docs.nosible.com/) and our [Python Package](https://github.com/NosibleAI/nosible-py). ``` from nosible import Nosible with Nosible(nosible_api_key="YOUR API KEY HERE") as nos: print(nos.search(prompt="Find technical blogs about Monte Carlo Tree Search")) ``` ## We trust Cybernaut-1 with our signals Cybernaut-1 is what we call a high-trust agentic search algorithm. What that means is that Cybernaut-1 has *direct access* to the internal logic in NOSIBLE. Our recent blog – ["The Road to Cybernaut-1: Rebuilding Search for AI"](https://nosible.com/blog/the-road-to-cybernaut-1) – goes into a lot of detail about what that internal logic encompasses. ![Cybernaut-1 concept diagram showing an agentic search system with access to NOSIBLE retrieval logic](https://nosible.com/images/2025/08/nosible-cybernaut-concept-2.png) Cybernaut-1 concept diagram showing an agentic search system with access to NOSIBLE retrieval logic We trust Cybernaut-1 with ALL the signals from NOSIBLE ## Cybernaut-1 uses them to self-improve Cybernaut-1 uses LLM-guided [Monte Carlo Tree Search](https://en.wikipedia.org/wiki/Monte_Carlo_tree_search) to iteratively construct a high-quality search that aligns with your given prompt. It balances exploration, exploitation, and inference cost by slowly moving from wide and shallow searches to narrow and deep searches. This approach is illustrated below: ![Cybernaut-1 search algorithm diagram showing LLM-guided Monte Carlo Tree Search from broad to focused queries](https://nosible.com/images/2025/08/nosible-cybernaut-algorithm.png) Cybernaut-1 search algorithm diagram showing LLM-guided Monte Carlo Tree Search from broad to focused queries Cybernaut-1 uses those signals to iterative and improve. ## So that you always get the best results Next week, we will be open-sourcing a comprehensive set of evaluations showing how Cybernaut-1 as well as Hybrid-3, the algorithm it uses, consistently match or outperforms leading search engines, even as we continue expanding our web coverage (currently growing at ~20 million webpages per day). In the meantime, we’d love for you to try it out. We are offering **$10,000 in Cybernaut-1 credits** to the first 20 AI startups that sign up early. [Building AI and looking for better context? Let's chat!](https://calendly.com/operations-nosible/) P.S. Looking for more of a technical deep dive? Go check out our blog: [*"The Road to Cybernaut-1: Rebuilding Search for AI*](https://nosible.com/blog/the-road-to-cybernaut-1) *".* Or, alternatively, check out: 1. [**Our Official Docs**](https://nosible-py.readthedocs.io/en/latest/) - how to use the Python package. 2. [**Our Python Package**](https://github.com/NosibleAI/nosible-py) - simply pip install nosible. 3. [**Our API documentation**](https://docs.nosible.com/) - For non-Python users. [All Research](https://nosible.com/blog) Related Research ![Railway-track illustration representing the road to Cybernaut-1](https://nosible.com/blog/illustrations/track.png) [Web Search](https://nosible.com/blog/tag/web-search) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) ### [The Road to Cybernaut-1: Rebuilding Search for AI](https://nosible.com/blog/the-road-to-cybernaut-1) 2025-08-20 17 min read ![High-resolution chart of selected 35B-A3B serving records across a 350-plus experiment campaign, from Qwen 3.5 FP8 on H200 at $0.218400 per million completion tokens in Experiment 006 to Qwen 3.6 source NVFP4 at $0.125959 in Experiment 332](https://nosible.com/images/2026/08/qwen-35b-a3b-serving-campaign-hero.svg) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) ### [What 350+ Experiments Taught Us About Serving Qwen 3.6 35B-A3B](https://nosible.com/blog/what-350-experiments-taught-us-about-serving-qwen-3-6) 2026-08-17 15 min read ![Thirty semantic factors rendered as daily time series on a dark grid, from geopolitical risk to firm-level political risk.](https://nosible.com/images/2026/08/semantic-factors-hero.png) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ### [The Future That Could Have Been: Turning Web Text into Semantic Stock Betas](https://nosible.com/blog/the-future-that-could-have-been-turning-web-text-into-semantic-stock-betas) 2026-08-05 11 min read > Meet Cybernaut-1, NOSIBLE’s agentic search system combining hybrid retrieval with LLM-guided Monte Carlo Tree Search. **URL:** https://nosible.com/blog/introducing-cybernaut-1-agentic-search-with-mcts --- --- title: "The Road to Cybernaut-1: Rebuilding Search for AI" description: "Why AI needs a purpose-built search engine, and how NOSIBLE rebuilt hybrid retrieval on the road to its Cybernaut-1 agentic search system." url: "https://nosible.com/blog/the-road-to-cybernaut-1" --- [Web Search](https://nosible.com/blog/tag/web-search) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) # The Road to Cybernaut-1: Rebuilding Search for AI [Stuart Reid](https://www.linkedin.com/in/stuartgordonreid/) 2025-08-20 17 min read Copy as Markdown ![Railway-track illustration representing the road to Cybernaut-1](https://nosible.com/blog/illustrations/track.png) In the near future all searches will be done by AIs operating on our behalf. The problem is that search was not built for them. It was built for us - no, let's be honest - it was built for advertisers. The consequence is that AI agents are being kneecapped. We will fix that problem. We are Search, rebuilt for AI. In this post, I will take you through our eight-stage multilingual retrieval pipeline step-by-step. I will touch on a variety of topics including tokenization, lexical search, semantic search, rank fusion, and more. Finally, I will share an important company update regarding an upcoming algorithm - cybernaut-1. ## AI needs its own search engine Building systems involves making trade-offs. How those trade-offs are decided cannot be disentangled from who your end user is. AIs are not people. People are not AIs. So, a search engine built for people will necessarily make different trade-offs to a search engine that is built for AIs. Take for example: - **Spell Checking**: AIs don’t need spell-checking. "Did you mean" CPU cycles are better spent on detecting threats like [prompt injections](https://en.wikipedia.org/wiki/Prompt_injection). - **Query Length**: Human queries are short and vague. AI queries are long and unambiguous, which negates the need for personalization. - **Few-Shot Search**: AIs can effortlessly generate multiple queries, so search engines that support multi-query input are more valuable. - **Recall Maxxing**: People need precision; AIs need recall. Short attention spans (people) versus massive context windows (AIs). - **Complexity**: Human search engines must be idiot-proof. AI search engines must be the opposite: *genius-friendly*. Why? Because a future where every search is done by AI is a future where every search is done by a SQL expert who speaks 200 languages and knows everything. I could go on for days about this, but you get the point: search engines are not aligned with AIs and AIs are paying the price. So, we are building the world's most AI-aligned search engine. Let's dive into how it actually works. ## Today we are 250,000 search engines From the outside, NOSIBLE looks and behaves like one large search engine. On the inside, NOSIBLE is actually made up of a federation of 250,000 smaller search engines called shards. Each shard acts like a giant magnet, attracting similar texts. Collectively they house almost a billion webpages, billions of embeddings, and hundreds of billions of words of text. ![NOSIBLE index architecture diagram showing a federation of 250,000 search shards](https://nosible.com/images/2025/08/nosible-index-architecture-2.png) NOSIBLE index architecture diagram showing a federation of 250,000 search shards NOSIBLE is a federation of 250,000 independent search engines each capable of performing lexical retrieval, semantic retrieval, full-text retrieval, hybrid retrieval, and more complex tasks. If you opened up a shard you would find: - **Metadata -** Descriptions, classifications, and more. - **Full-Text Index -** Used for longer n-gram retrieval. - **Lexical Index (BM25) -** Used for [keyword-based retrieval](https://en.wikipedia.org/wiki/Okapi_BM25). - **Semantic Index -** Used for [vector-based retrieval](https://nosible.com/blog/ensemble-and-distil). - **Vocabulary -** The set of unique words in this shard. - **Keywords -** Important words to this shard. - **Entities -** Important entities to this shard. - **Synonyms -** A graph learned from word co-occurrences. - **Classifiers -** Fine-tuned classifiers for signals e.g. sentiment. - **Compressors -** Trained compression dictionaries. - **Bloom Filters -** Used to quickly check if phrases exist. - **Queries -** Examples of queries that align with this shard. - **Evals -** Gold standard datasets for evaluating this shard. - **Rerankers -** Fine-tuned neural reranking models. - **Write-Ahead Log -** Pending entries to the shard. And hopefully soon you would also find: - **LoRAs -** Shard [Low-Rank Adaptations](https://arxiv.org/abs/2106.09685). - **Small LLMs -** 0.5-2Bn parameter LLMs (more later). Each shard is independently capable of performing lexical retrieval, semantic retrieval, full-text retrieval, hybrid retrieval, and more complex tasks. Each shard can be deployed on its own or alongside others. This gives us a lot of options for scaling up to 10+ billion webpages within the next 6 months. Shards are learned using an efficient algorithm we developed for identifying semantically and lexically coherent clusters of texts that are also approximately uniformly sized. Uniformity helps ensure that queries per second per shard is predictable. Our most recent training run involved 250,000,000 embeddings. Our next run will involve 1 billion embeddings and a *lot* of GPU time. Let's take shard 11,343 as an example: > **11,343: CRISPR Genome Editing Using sgRNA, Plasmids, Nucleases, and Homologous Recombination in Bacterial and Yeast Systems** 11,343 focuses on CRISPR-based genome editing methods involving sgRNA, crRNA, and nucleases to induce targeted DNA cleavage. It covers plasmid design, transfection, and the use of homologous recombination and non-homologous end joining for mutation and gene insertion. The content includes molecular tools like PCR, cloning, and protein expression in bacterial and yeast models. It addresses spacer sequences, codon optimization, and biosynthetic pathways, highlighting applications in mutagenesis, transgenesis, and metabolic engineering. Key elements include DNA repair mechanisms, vector construction, and functional assays to study gene function and phenotype changes. - **Classifications:** - Language: Mostly English - Geography: Worldwide - Brand Safety Classification: Safe - Industry Classification: Biotechnology - IAB: Biotech and Biomedical Industry - **Scope & Size:** - Number of documents: 134,523 - Number of words: 8,446,510 - Vocabulary size: 64,992 - Number of keywords: 2,523 - **Top 20 Keywords:** - plasmid → plasmid, plasmids - crispr → CRISPR, CRISPR-Cas9, CRISPR-Cas systems - sgrna → single guide RNA, sgRNA - genom → genome, genomic, genomes - grna → guide RNA, gRNA - rna → RNA - nucleas → nuclease, nucleases (e.g., Cas9 nuclease) - crrna → CRISPR RNA, crRNA - sgrnas → single guide RNAs, sgRNAs - transfect → transfect, transfection, transfected - nucleotid → nucleotide, nucleotides - grnas → guide RNAs, gRNAs - streptomyc → Streptomyces (bacterial genus often used in biotechnology) - rnas → RNAs (plural of RNA, could be mRNAs, tRNAs, etc.) - spacer → spacer sequence(s), CRISPR spacer - codon → codon, codons, codon optimization - peptid → peptide, peptides - strep → Streptococcus - gfp → green fluorescent protein (GFP) - recombin → recombination, recombinant, recombine - **Top 20 Entities** - [CRISPR Therapeutics](https://en.wikipedia.org/wiki/CRISPR_Therapeutics) - Mentioned 6408 times - [Thermo Fisher Scientific Inc.](https://en.wikipedia.org/wiki/Thermo_Fisher_Scientific) - Mentioned 2027 times - [Addgene](https://en.wikipedia.org/wiki/Addgene) - Mentioned 1551 times - [Qiagen](https://en.wikipedia.org/wiki/Qiagen) - Mentioned 1374 times - [Illumina](https://en.wikipedia.org/wiki/Illumina,_Inc.) - Mentioned 1204 times - [Invitrogen](https://en.wikipedia.org/wiki/Invitrogen) - Mentioned 1160 times - [New England Biolabs](https://en.wikipedia.org/wiki/New_England_Biolabs) - Mentioned 1137 times - [Editas Medicine](https://en.wikipedia.org/wiki/Editas_Medicine) - Mentioned 1014 times - [British Gynaecological Cancer Society](https://en.wikipedia.org/wiki/British_Gynaecological_Cancer_Society) - Mentioned 938 times - [Intellia Therapeutics](https://en.wikipedia.org/wiki/Intellia_Therapeutics) - Mentioned 884 times - [Life Technologies](https://en.wikipedia.org/wiki/Life_Technologies) - Mentioned 800 times - [Pacific Biosciences](https://en.wikipedia.org/wiki/Pacific_Biosciences) - Mentioned 648 times - [Takara Bio](https://en.wikipedia.org/wiki/Takara_Holdings) - Mentioned 565 times - [Agilent Technologies](https://en.wikipedia.org/wiki/Agilent_Technologies) - Mentioned 558 times - [Mammoth Biosciences](https://en.wikipedia.org/wiki/Mammoth_Biosciences) - Mentioned 529 times - [Sigma-Aldrich](https://en.wikipedia.org/wiki/Sigma-Aldrich) - Mentioned 496 times - [Promega](https://en.wikipedia.org/wiki/Promega) - Mentioned 437 times - [GenBank](https://en.wikipedia.org/wiki/GenBank) - Mentioned 421 times - [Integrated DNA Technologies](https://en.wikipedia.org/wiki/Integrated_DNA_Technologies) - Mentioned 401 times - [Millipore Corporation](https://en.wikipedia.org/wiki/Merck_Millipore) - Mentioned 397 times And that's just one! We have 250,000 of them and collectively they span every aspect of the human condition past and present. They also provide a unique glimpse into how models perceive / misperceive our world. ## The journey from questions to answers Every question that comes into NOSIBLE goes through a pretty sophisticated eight-stage retrieval pipeline. These eight stages are as follows: 1. **Language Detection and Translation** 2. **Multilingual Tokenization** 3. **Search Intent Prediction** 4. **Instruction Tuning and Embedding** 5. **Shard Selection** 6. **Shard Reranking** 7. **Shard-based Question Expansion** 8. **Retrieval using Map-Reduce** Let's follow *"What lessons from bacteria and yeast actually translate into safer gene-editing medicines?"* through the pipeline in English and in Japanese. ### **1 // Language Detection and Translation** Due to the ever shrinking - but unfortunately still present - [language gap](https://jina.ai/news/bridging-language-gaps-in-multilingual-embeddings-via-contrastive-learning/) in multilingual embedding models it is still necessary to detect what language the question and expansions were written in and, if results are requested in a different language, translate that question into the requested language. **Worked Example** - "What lessons from bacteria and yeast actually translate into safer gene-editing medicines?" → "What lessons from bacteria and yeast actually translate into safer gene-editing medicines?" (no translation). - "What lessons from bacteria and yeast actually translate into safer gene-editing medicines?" → "細菌と酵母から得られる教訓は、より安全な遺伝子編集医薬品にどのように応用されるのでしょうか?". **Tech Stack** - [fasttext-langdetect](https://github.com/zafercavdar/fasttext-langdetect) - supports 170+ languages. - [gemma-3n-e4b-it](https://openrouter.ai/google/gemma-3n-e4b-it) - supports 140 languages. ### **2 // Multilingual Tokenization** Next, the question and expansions are tokenized. This involves [sentence boundary detection](https://en.wikipedia.org/wiki/Sentence_boundary_disambiguation), [text segmentation](https://en.wikipedia.org/wiki/Text_segmentation), [text inflection](https://en.wikipedia.org/wiki/Inflection), [Unicode normalization](https://en.wikipedia.org/wiki/Unicode_equivalence#Normalization), [word stemming](https://en.wikipedia.org/wiki/Stemming), [stop-word removal](https://en.wikipedia.org/wiki/Stop_word), and other classic NLP techniques. Getting this to work equally well across many languages is difficult. We ended up creating a standardized interface that wraps probably a dozen or more different NLP packages and abstracts away all language-specific complexities. I know it might sound strange to be working on this in 2025. Why not just use an LLM? Why not just do semantic search? The answers to those questions are (1) even small LLMs are too slow and (2) semantic search is not a silver bullet. In all honesty, well-calibrated lexical search indices beat well-calibrated semantic search indices on AI queries more often than not. Why? Because AI queries are longer and have high intent. That said, [hybrid search](https://weaviate.io/blog/hybrid-search-explained) dominates both. **Worked Example** Q: "What lessons from bacteria and yeast actually translate into safer gene-editing medicines?" yields the following tokens: - "yeast" - "bacteria" - "safer" - "lesson" - "translat" - "gene" - "medicin" - "edit" - "actual" Q: "細菌と酵母から得られる教訓は、より安全な遺伝子編集医薬品にどのように応用されるのでしょうか?" yields the following tokens: - "細菌" (bacteria) - "教訓" (lesson) - "遺伝子" (gene) - "酵母" (yeast) - "応用" (application) - "治療" (treatment) - "編集" (edit) - "どの" (which) - "安全" (safety) **Tech Stack** - [SpaCy](https://github.com/explosion/spaCy) - full/partial support for 75 languages - [NLTK](https://github.com/nltk/nltk) - mixed support for ~50 languages. - [BlingFire](https://github.com/microsoft/BlingFire) - English only (super-fast) - [pySBD](https://github.com/nipunsadvilkar/pySBD) - support for 22 languages - [pyStemmer](https://github.com/snowballstem/pystemmer) - support for 24 languages - [indic-NLP-library](https://github.com/anoopkunchukuttan/indic_nlp_library) - supports Indian languages - [pymorphy3](https://github.com/no-plagiarism/pymorphy3) - support for Russian and Ukrainian. - [jieba](https://github.com/fxsjy/jieba) - support for Chinese text. - [mecab-python3](https://github.com/SamuraiT/mecab-python3) - support for Japanese text. - [mecab-ko](https://github.com/shirakaba/mecab-ko) - support for Korean text. - [wtpsplit](https://github.com/segment-any-text/wtpsplit) - 85 languages (incredibly slow). - [ipadic](https://github.com/polm/ipadic-py) - support for Japanese text. - [inflect](https://github.com/jaraco/inflect) - support for English text only. - regex - sigh. Way too much regex. ### **3 // Search Intent Prediction** Next, we run the tokens and their [TF-IDF scores](https://en.wikipedia.org/wiki/Tf%E2%80%93idf) through a system that identifies search intents. We define a search intent as a sequence of proximal tokens that have a high " [harmonic TF-IDF score](https://en.wikipedia.org/wiki/Harmonic_mean) " and do not contain certain [parts-of-speech](https://en.wikipedia.org/wiki/Part-of-speech_tagging) like, for example, pronouns, determiners, conjunctions, adverbs, etc. **Worked Example** Q: "What lessons from bacteria and yeast actually translate into safer gene-editing medicines?" yields the following search intents: - "editing medicine" - "gene-editing" - "gene-editing medicine" - "safer gene" - "safer gene-editing" - "safer gene-editing medicine" Currently search intents are not available in Japanese. We expect to have search intents working for all supported languages in the week. **Tech Stack** - [NumPy](https://github.com/numpy/numpy) - [SpaCy](https://github.com/explosion/spaCy) ### **4 // Instruction Tuning and Embedding** Next, we generate an appropriate instruction for the question and submit it, along with the expansions, to [multilingual-e5-large-instruct](https://huggingface.co/intfloat/multilingual-e5-large-instruct). E5 is an open-source, instruction-tuned embedding model from Microsoft Research. It is designed to align vectors with natural language instructions. A while back we used an LLM to generate and "evolve" optimal instruction templates for E5. Here is one of the most successful templates: > "Given a question, please retrieve any relevant {Language} Headlines, Leads, Passages, and Source URLs that focus on the same named entities as the question, and provide substantive answers to the question." We have templates that have placeholders for named entities, geographic regions, topics of interest, and other dimensions. In our internal evals, we have seen that optimizing the instruction - or getting a small LLM to write a bespoke one - can yield a free 1-5% improvement in search precision and recall. **Tech Stack** - [Hugging Face](https://huggingface.co/) ### **5 // Shard Selection** Next, we move on to selecting the best shards for the question. This is arguably the single most important stage in our retrieval pipeline because, if we route the question to the wrong shards, we won't return the best document. Our current approach involves computing four ranking factors, namely: - **Vanilla Dense Similarity -** This selector computes the cosine similarity between the E5 embedding of the question and E5 embeddings of the LLM-generated summaries for each shard ([uses HNSW](https://en.wikipedia.org/wiki/Hierarchical_navigable_small_world)). - **Bayesian Dense Similarity -** This selector is derived from a vector search index we developed that often beats the current state-of-the-art. I can't discuss it because we are pursuing a patent on it 😉. - **Vanilla Sparse Similarity -** This selector computes the sparse similarity between the keywords in the question and the keywords in each shard. It's essentially just a massive sparse TF-IDF matrix. - **Entity Sparse Similarity -** If the question contains one or more entities, we also compute the sparse similarity between those entities and the entities in each shard. Again, it is a massive sparse matrix. These ranking factors are then combined using [reciprocal rank fusion](https://medium.com/@devalshah1619/mathematical-intuition-behind-reciprocal-rank-fusion-rrf-explained-in-2-mins-002df0cc5e2a) (RRF). RRF is a simple but powerful ensembling method that merges multiple ranked lists by giving higher weight to items that appear near the top across rankings. This gives robust results even if one of the ranking factors is noisy. Because our shards are trained to be maximally semantically and lexically coherent, we are able to select the best shards almost all of the time. The only time we struggle is for very short queries. But, like I said, we are building NOSIBLE for AIs, not people, so this is a trade-off we are happy to make. **Worked Example** Q: "What lessons from bacteria and yeast actually translate into safer gene-editing medicines?" is routed to these English shards: - **220,067** - Title: CRISPR Genome Editing RNA Enzymes Plasmids Epigenetics Molecular Biology Protein Mutation Recombinant Technologies - **49,619** - Title: Overview of CRISPR Genome Editing Technologies and Molecular Mechanisms in Genetic Engineering and Therapeutics - **178,396** - Title: Overview of GMO Safety, Toxicology, Biotechnology, Regulatory Agencies, and Food Additives Impacting Agriculture and Health - **+97 others** Q: "細菌と酵母から得られる教訓は、より安全な遺伝子編集医薬品にどのように応用されるのでしょうか?" is routed to these Japanese shards: - **199,016** - Title: 量子コンピューティング技術開発と応用に関する研究とシステム設計の最新動向と課題解決方法 - **220,938** - Title: セキュリティとサイバーセキュリティに関するデータ保護システム開発と企業向け安全管理ソリューションの概要 - **111,375** - Title: 企業のデジタルサービス開発とプラットフォーム活用によるビジネス成長と技術課題解決の方法論 - **+97 others** **Tech Stack** - [NumPy](https://github.com/numpy/numpy) - [SciPy](https://github.com/scipy/scipy) - [Numba](https://github.com/numba/numba) - [SimSIMD](https://github.com/ashvardanian/simsimd) - [USearch](https://github.com/unum-cloud/usearch) ### **6 // Shard Reranking** In the previous step the shards we selected for the Japanese question were quite bad. The first shard is about quantum computing, the second shard is about data protection, and the third shard is about digital platforms. This is where shard [reranking](https://aiwiki.ai/wiki/Re-ranking) comes in. Reranking involves reordering the selected shards by estimating their actual relevance to the question, so that the best shards rise to the top and irrelevant ones (like the ones we selected) get pushed down. We have quite a few rerankers: - **Bloom Filter Reranker -** Each shard has a [bloom filter](https://en.wikipedia.org/wiki/Bloom_filter) that keeps track of important phrases it contains. If a selected shard hasn't seen any of the user's search intents, then that shard should get downranked. - **Compression Reranker -** Each shard has a trained [Zstandard text compression](https://proceedings.mlr.press/v163/kasturi22a) dictionary. The shard that can compress the input question the best is more likely to contain the most similar content. - **Neural Reranker -** We can also feed the LLM-generated titles and descriptions of each selected shard along with the question into a neural reranker like [bge-reranker-v2-m3](https://huggingface.co/BAAI/bge-reranker-v2-m3) to get a reranked set of shards. - **Page (Re)Ranker -** Last, but certainly not least, it is also possible to use a nearest neighbor graph and the [Personalized PageRank](https://cs.stanford.edu/people/plofgren/Fast-PPR_KDD_Talk.pdf) algorithm to produce a probability mass over your selected shards. Once again, these ranking factors are then combined using reciprocal rank fusion (RRF). We have also used LLMs to rerank selected shards. That works very well but the added latency is simply too high for a search engine. **Worked Example** In the case of the Japanese shards the one and only shard we have that talks about gene editing in Japanese bubbles to the top. In the case of the English shards the shard that actually talks about bacteria and yeast in the context of gene editing bubbles up to the third position. Which is what we want to see. Q: "What lessons from bacteria and yeast actually translate into safer gene-editing medicines?" is routed to these English shards: - **49,619** - Title: Overview of CRISPR Genome Editing Technologies and Molecular Mechanisms in Genetic Engineering and Therapeutics - **220,067** - Title: CRISPR Genome Editing RNA Enzymes Plasmids Epigenetics Molecular Biology Protein Mutation Recombinant Technologies - **11,343** - Title: CRISPR Genome Editing Techniques Using sgRNA, Plasmids, Nucleases, and Homologous Recombination in Bacterial and Yeast Systems - **+97 others** Q: "細菌と酵母から得られる教訓は、より安全な遺伝子編集医薬品にどのように応用されるのでしょうか?" is routed to these Japanese shards: - **200,739** - Title: 医薬品開発と臨床試験における製薬企業の治療法承認市場展開と感染症ワクチン技術動向分析 - **30,173** - Title: 睡眠の質向上に関する研究と健康影響 食事習慣 ホルモンバランス ストレス対策 運動効果 栄養摂取の重要性 - **168,764** - Title: ストレス管理と健康維持に関する食事、栄養素、メンタルヘルス、育児、免疫、生活習慣の重要性について - **+97 others** The translation of shard 200,739 in English is *"Analysis of Pharmaceutical Companies’ Therapeutic Approval, Market Deployment, and Infectious Disease Vaccine Technology Trends in Drug Development and Clinical Trials".* **Tech Stack** - [Zstandard](https://github.com/facebook/zstd) - [rBloom](https://github.com/KenanHanke/rbloom) ### **7 // Shard-based Question Expansion** Next, we use the synonym graphs in each shard to probabilistically expand our search terms. Because each shard is lexically and semantically coherent, the synonyms in each shard are unambiguous. Put simply, "gene" in shard 11,343 only has the genetic meaning. We don't need to worry that Gene is also a name and, in some other shards, would co-occur a lot with "Willy Wonka". It's worth mentioning that this step can introduce some noise / serendipity, so we are very careful not to overwhelm the original search words. **Worked Example** Q: "What lessons from bacteria and yeast actually translate into safer gene-editing medicines?" when expanded produces the following: - Original Search Words: - "yeast" - "bacteria" - "safer" - "lesson" - "translat" - "gene" - "medicin" - "edit" - "actual" - Expanded Search Words - **"yeast"** - **"bacteria"** - **"safer"** - **"lesson"** - **"translat"** - **"gene"** - **"medicin"** - **"edit"** - **"actual"** - "invad" [NEW] - "palindrom" [NEW] - "bacteriophag" [NEW] - "interspac" [NEW] - "antibiot" [NEW] - "archaea" [NEW] - "bacterium" [NEW] - "pathogen" [NEW] - "phage" [NEW] - "dextros" [NEW] - "infecti" [NEW] - "raffinos" [NEW] - "biofuel" [NEW] - "galactos" [NEW] - "acronym" [NEW] - "chromosom" [NEW] - "crispr" [NEW] - "insect" [NEW] - "prokaryot" [NEW] - "mojica" [NEW] - "bacto" [NEW] Currently question expansions are not available in Japanese. We expect to have question expansions working for all supported languages in the week. **Tech Stack** - [NumPy](https://github.com/numpy/numpy) ### **8 // Retrieval Using Map-Reduce** At this point we finally have (1) our instruction-optimized embeddings, (2) our expanded set of search terms, (3) our search intents, and (4) a list of 10-30 shards that almost certainly contain documents that relate directly to our question. This package of information is then broadcast to each of the final shards. When received each shard will: 1. Run the user's SQL filter to filter out documents. 2. Execute a full lexical search over the valid records. 3. Execute a full semantic search over the valid records. 4. Fuse the lexical and semantic results together. 5. Then do a full-text search over top results for intents. The results from each shard are combined and grouped by document identifier. Once the search completes a highly relevant snippet is constructed for each search result. The maximum length of the snippet is decided by the user. **Worked Example** Q: "What lessons from bacteria and yeast actually translate into safer gene-editing medicines?" yields the following search results: **🔗** [**Yeast could prove game-changer in medicine**](https://www.euronews.com/next/2019/03/25/yeast-could-prove-game-changer-in-medicine) *In this project, we use genes sourced from plants to make a metabolite that is produced by plants such as oranges — we're producing naringenin, a molecule produced by fruit trees like lemons, grapefruits or oranges," said Molecular microbiologist Jean-Marc Daran of the Delft University of Technology. **To reprogram yeast, researchers at the Delft University of Technology are using an editing process called CRISPR. The procedure inserts genes from plants or bacteria into yeast. That can change how the cell factories work — and even how they smell!** "We edited the DNA, and we've added among others one gene from a plant. And because of that it now also really smells like roses. Your lab smells better of course, but we are more looking into the industrial applications so we can smell this as a flavour and aroma compound — mostly for the cosmetic industry, mostly perfumes but also mascara, lipstick — they all have this flavour compound in them," metabolic engineering researcher Jasmijn Hassing told Euronews.* **+99 more results** Q: "細菌と酵母から得られる教訓は、より安全な遺伝子編集医薬品にどのように応用されるのでしょうか?" yields the following search results: [**🔗 あらゆる産業の常識を変えうる「エンジニアリングバイオロジー」が期待されるわけ**](https://thebridge.jp/2025/02/why-engineering-biology-is-expected-to-be-the-game-changer-in-every-industry-gb-universe) *エンジニアリングバイオロジーとは、 バイオテクノロジーを用いて有用な機能を持った生物(スマートセル)を開発し、バイオプロセスによって有用物質を生産する技術領域 のことを指します。 スマートセルとは、バイオテクノロジーによって有用物質の生産能力を向上させた、細菌や酵母、植物などを総称した日本での呼び名です(海外では「Engineered organisms」や「Engineered cells」など)。 スマートセルを活用したバイオプロセスによるバイオものづくりは 医薬品、食料、燃料、衣類など幅広い分野への応用ができます 。* *Engineering biology refers to a technological domain in which organisms with useful functions (“smart cells”) are developed using biotechnology, and useful substances are then produced through bioprocesses. The Japanese term “smart cells” collectively refers to bacteria, yeast, plants, and other organisms whose ability to produce useful substances has been enhanced through biotechnology (overseas, they are called things like “engineered organisms” or “engineered cells”). Bio-manufacturing based on smart cells and bioprocesses can be applied across a wide range of fields, including pharmaceuticals, food, fuel, and clothing.* **+99 more results** **Tech Stack** - [Polars](https://github.com/pola-rs/polars) - [SimSIMD](https://github.com/ashvardanian/simsimd) - [Numba](https://github.com/numba/numba) - [NumPy](https://github.com/numpy/numpy) - [ahocorasick-rs](https://github.com/G-Research/ahocorasick_rs) ## How does this tie into agentic search? As I mentioned earlier - Human search engines must be idiot-proof; AI search engines must be the opposite: *genius-friendly*. In the coming days we will launch our V2 search API which includes an algorithm called cybernaut-1. > Cybernaut: A voyager in cyberspace, indicating someone who navigates or explores online environments. Cybernaut-1 is an AI agent with unrestricted access to *everything* in NOSIBLE including every shard, algorithm, selector, reranker, and signal. It knows what these things are and can tune them on the fly to find better results. Other agentic searchers, by comparison, are flying blind. They only see the top search results. They don't know how they got there. So, their ability to reflect and iterate is fundamentally constrained. It's just a random walk. Put differently, cybernaut-1 is a reinforcement learner that designs and iterates on search policies to maximize search relevancy. It is the culmination of months of work, and it will be generally available on the 25th of August 2025. ![Cybernaut-1 concept diagram showing reinforcement learning used to refine search policies](https://nosible.com/images/2025/08/nosible-cybernaut-concept-1.png) Cybernaut-1 concept diagram showing reinforcement learning used to refine search policies The relationship between cybernaut-1 and NOSIBLE is a high-trust one. We trust that cybernaut-1 is able to understand and navigate the complexity of our search engine, so we surface it instead of hiding it. --- ## Epilogue. At NOSIBLE we believe that search is a scaling law on par data or compute. Giving AIs access to 10x, 100x, 1000x more search will make them more intelligent and more capable. [Recent research has hinted that this is true.](https://arxiv.org/abs/2410.04343) We believe this because it reflects how we work. Our intelligence is not an island. Nobody expects us to recall every fact from memory. We are deeply connected to knowledge 24/7. We seek it, read it, grok it, and act upon it. AIs will do the same, only at superhuman scale. AIs will seek knowledge millions of times a day in ways that are markedly different from you or I. So, they need their own search engine. Search isn't a tool call - it is intelligence. [All Research](https://nosible.com/blog) Related Research ![Cybernaut-1 illustration representing agentic search with Monte Carlo Tree Search](https://nosible.com/blog/illustrations/cyber.png) [Web Search](https://nosible.com/blog/tag/web-search) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) ### [Introducing Cybernaut-1: Agentic Search using MCTS](https://nosible.com/blog/introducing-cybernaut-1-agentic-search-with-mcts) 2025-08-26 2 min read ![High-resolution chart of selected 35B-A3B serving records across a 350-plus experiment campaign, from Qwen 3.5 FP8 on H200 at $0.218400 per million completion tokens in Experiment 006 to Qwen 3.6 source NVFP4 at $0.125959 in Experiment 332](https://nosible.com/images/2026/08/qwen-35b-a3b-serving-campaign-hero.svg) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) ### [What 350+ Experiments Taught Us About Serving Qwen 3.6 35B-A3B](https://nosible.com/blog/what-350-experiments-taught-us-about-serving-qwen-3-6) 2026-08-17 15 min read ![Thirty semantic factors rendered as daily time series on a dark grid, from geopolitical risk to firm-level political risk.](https://nosible.com/images/2026/08/semantic-factors-hero.png) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ### [The Future That Could Have Been: Turning Web Text into Semantic Stock Betas](https://nosible.com/blog/the-future-that-could-have-been-turning-web-text-into-semantic-stock-betas) 2026-08-05 11 min read > Why AI needs a purpose-built search engine, and how NOSIBLE rebuilt hybrid retrieval on the road to its Cybernaut-1 agentic search system. **URL:** https://nosible.com/blog/the-road-to-cybernaut-1 --- --- title: "A Pattern for Scaling the Value Proposition of LLMs: Ensemble and Distil 🚀" description: "Train a simple regression on sentence embeddings to distil an LLM ensemble—and outperform GPT-4 on financial sentiment classification." url: "https://nosible.com/blog/ensemble-and-distil" --- [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [Financial Sentiment](https://nosible.com/blog/tag/financial-sentiment) # A Pattern for Scaling the Value Proposition of LLMs: Ensemble and Distil 🚀 [Stuart Reid](https://www.linkedin.com/in/stuartgordonreid/) 2024-02-06 12 min read Copy as Markdown ![3D cube illustration representing LLM ensemble distillation](https://nosible.com/blog/illustrations/cube.png) In our [previous post](https://nosible.com/blog/news-sentiment-showdown-who-checks-vibes-best) we showed that Large Language Models (LLM) including PaLM-2, Gemini, and GPT-4 are able to assign sentiment labels with exceptional accuracy. We also showed that they massively outperform previous state of the art models like FinBERT. But it wasn't all good news. Using LLMs to label a corpus of 10 million news stories would cost thousands of dollars, take weeks to finish, and would be nearly impossible to scale to near-real-time use cases. Which brings us to the inevitable question that goes through every person's mind when they attempt to scale the capabilities of Large Language Models... > **Well, sh*t! Now what?** Fortunately, there is a widely used pattern in data science that can help. I like to call it "ensemble and distil". Here's how it works in simple terms: 1. **Curate** a dataset representative of your whole corpus. 2. **Label** that dataset using a mix of large language models. 3. **Ensemble** the models to create a teacher model. 4. **Distil** the capability of the teacher into a student model. 5. **Scale** the student model to the rest of your corpus. In this post we will show you *how* this is done. But, more importantly, we will explain why this *ought* to be done. At the risk of spoiling the surprise, we will also show that a simple linear regression trained to emulate the outputs of an LLM-ensemble from off-the-shelf sentence transformer embeddings can outperform GPT-4 at sentiment classification! Yeah, you read that right. *Before we get into the details, I want to use this opportunity to say a public thank you to* [*Moabi Mokhoro*](https://www.linkedin.com/in/moabi-mokhoro/) *and* [*Deon van Heerden*](https://www.linkedin.com/in/deon-van-heerden/) *. During their machine learning internship at Nosible they worked on this project and taught us a lot.* ## **Background** In the [previous post](https://nosible.com/blog/news-sentiment-showdown-who-checks-vibes-best) we discussed the first and second step. If you have the time, I recommend reading it for added context. But here's a quick recap. We curated a dataset of 10,378 [news stories from our corpus](https://nosible.com/blog/using-vector-search-to-see-signals-in-company-news). This dataset has news stories from all years, all industries, and all news categories. Next, we evaluated a variety of sentiment classification models on this dataset. These included: TextBlob, VADER, Flair, SigmaFSA, FinBERT, FinBERT-Tone, Bison, Unicorn, Gemini-Pro, GPT-3.5-Turbo, GPT-4, and GPT-4-Turbo. Using human agreement as our measure of success we observed that the top three models were: Unicorn (84%), GPT-4 (74%), and Gemini-Pro (67%). ## **The Teacher** The goal of [ensembling](https://en.wikipedia.org/wiki/Ensemble_learning) is to combine multiple models to create a new model that is more capable than any model individually. This does not always happen. More often than not, an ensemble of models is likely to produce a *mediocre* result. For the ensemble to be better than the best model, the stars need to align. First, the individual models must be better than random. Second, the models must be correlated when they are right. And third, the models must be uncorrelated when they are wrong. If, and only if, these stars align is there a chance that an ensemble will be better than the best. Techniques like boosting, which train models on the errors of other models, do this explicitly. One trick I like to use when building ensembles is iterative addition. This is a greedy procedure that starts off with the best model and then iteratively includes the most "additive" model until no more models are additive. Additive models are ones that improve the accuracy of the ensemble at each point. When applied to the sentiment models this is what we got: - Start with `Unicorn` (83.6%) - Then add `GPT-3.5` (+2.80% boost) - Then add `GPT-4` (+2.00% boost) - Then add `SigmaFSA` (+1.60% boost) - Then add `Bison` (+0.40% boost) The resulting ensemble ends up with an agreement ratio of **90.4%** with human labels. That is 6% better than the best model (Unicorn) in absolute terms. Because iterative addition is extremely quick, you can repeat this multiple times over a sample of your data to arrive at a probability distribution over the possible ensembles. Doing this is a useful test for overfitting risk. Across 1,000 simulations this ensemble was the best 51% of the time. The next best ensemble was the same except that it excluded Bison. That was the best 13.8% of the time. And, finally, the third best ensemble was the same except that it included TextBlob-0.30. That was the best 11% of the time. Which is all just to say, we're confident that this is, in fact, the best ensemble. ![Comparison of candidate LLM ensembles by the share of simulations in which each produces the best result](https://nosible.com/images/2024/01/ensemble_model.png) Comparison of candidate LLM ensembles by the share of simulations in which each produces the best result The best ensemble found through iterative addition beats the best model we tested and agrees with the human-assigned labels a stunning 90% of the time. There you have it, our teacher will be an ensemble that aggregates the outputs of Unicorn, GPT-3.5, GPT-4, SigmaFSA, and Bison. In order to get discrete labels from this ensemble we threshold it such that any aggregates greater than 1 are "Positive". Any aggregates less than -1 are "Negative". And anything in between is "Neutral". The goal is now to get our students to emulate the teacher. ## **The Students** To keep things nice and simple, all of the students in our experiment are simple [ordinary least squares regression](https://en.wikipedia.org/wiki/Ordinary_least_squares) models. Over the years I've found that a linear regression fitted on good features is always hard to beat. Linear regression just works. It's also efficient, portable, and generalizes well. The only difference between each student was the input data they were fitted on. For our students to learn from the teacher, we need to give them something to learn from. In the spirit of keeping with the student-teacher analogy, let's call that the "textbook". In 2024 the best textbook for our students to learn from are sentence embeddings from pretrained transformer models. In the process of learning how to solve a natural language problem, neural networks transform the sentences they are given into sets of numbers. These sets of numbers are called [embeddings or vectors](https://en.wikipedia.org/wiki/Sentence_embedding), and they are the network's internal representation of the sentences. Embedding is a superpower of neural networks because it allows us to turn variable length sentences into fixed length vectors that capture a great deal of the information contained in the sentence. The biggest challenge now is finding the right embedding model. Every week more models appear on the [Massive Text Embedding Benchmark Leaderboard](https://huggingface.co/spaces/mteb/leaderboard) making it harder to navigate. Luckily for you, we look at it often and had a good idea which sentence transformers we wanted to evaluate for this blog post. ## **The Results** We ended up with 38 models in this survey. These included the teacher, the 7 LLMs we [evaluated last week](https://nosible.com/blog/news-sentiment-showdown-who-checks-vibes-best), and 30 ordinary least squares regression students that were trained to emulate the outputs of the teacher using various sentence transformer embeddings. The student models are denoted by `OLS(name-of-sentence-transformer)` to make it obvious what's what. One unsurprising result was that the students fitted on embeddings from larger sentence transformer models did better than the students fitted on embeddings from smaller sentence transformer models. More surprisingly, when we looked at the data, we noticed that the E5 models from Microsoft were consistently "punching about their weight" in every class. Here's what I mean: - Best with <100Mn parameters? `e5-small-v2` - Best with 100Mn to 200Mn parameters? `e5-base-v2` - Best with 300Mn to 400Mn parameters? `e5-large-v2` ![Benchmark chart showing E5 embedding models outperforming larger language models on financial sentiment](https://nosible.com/images/2024/02/e5-models-punching-up.png) Benchmark chart showing E5 embedding models outperforming larger language models on financial sentiment Even more surprisingly were how well the *unsupervised* versions of the E5 sentence transformers performed. We thought there would be a big difference between the two versions, but the differences were almost insignificant. Going forward we will be using the `multilingual-e5-large` embeddings. For more information on these models, please see this paper from Microsoft: > [**"Text Embeddings by Weakly-Supervised Contrastive Pre-training"**](https://huggingface.co/papers/2212.03533) by Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei. Because that's quite a lot of models we've broken up the results into the following categories - models with 384-dimensional vectors, models with 768-dimensional vectors, models with 1024-dimensional vectors, and multilingual models with either 384, 768, or 1024-dimensional vectors. Each result table shows the type of model, the number of likes on Hugging Face, the number of downloads on Hugging Face, and the **out-of-sample accuracy** versus the Teacher model. ### **384-Dimensional Vectors** ![Financial-sentiment benchmark results for 384-dimensional sentence embeddings](https://nosible.com/images/2024/02/384-dimensional.png) Financial-sentiment benchmark results for 384-dimensional sentence embeddings - **LLM Targets:** - PaLM-2 Unicorn (89.04%) - GPT 4 Original (79.90%) - **Students:** - Best: `e5-small-v2` (76.54%) - Fastest: `all-MiniLM-L6-V2` (68.13%) - Most Popular: `all-MiniLM-L6-V2` (68.13%) ### **768-Dimensional Vectors** ![Financial-sentiment benchmark results for 768-dimensional sentence embeddings](https://nosible.com/images/2024/02/768-dimensional.png) Financial-sentiment benchmark results for 768-dimensional sentence embeddings - **LLM Targets:** - PaLM-2 Unicorn (89.04%) - GPT 4 Original (79.90%) - **Students:** - Best: `sentence-t5-large` (80.40%) 🏆 - Fastest: `all-distilroberta-V1` (73.34%) - Most Popular: `all-mpnet-base-v2` (75.46%) ### **1024-Dimensional Vectors** ![Financial-sentiment benchmark results for 1024-dimensional sentence embeddings](https://nosible.com/images/2024/02/1024-dimensional.png) Financial-sentiment benchmark results for 1024-dimensional sentence embeddings - **LLM Targets:** - PaLM-2 Unicorn (89.04%) - GPT 4 Original (79.90%) - **Students:** - Best: `e5-large-v2` (80.21%) 🏆 - Fastest: `bge-large-en-v1.5` (76.58%) - Most Popular: `bge-large-en-v1.5` (76.58%) ### **Multilingual Vectors** ![Financial-sentiment benchmark results for multilingual sentence embeddings](https://nosible.com/images/2024/02/multilingual.png) Financial-sentiment benchmark results for multilingual sentence embeddings - **LLM Targets:** - PaLM-2 Unicorn (89.04%) - GPT 4 Original (79.90%) - **Students:** - Best: `paraphrase-multilingual-base-v2` (78.63%) - Fastest: `multilingual-e5-small` (72.99%) - Most Popular: `distil-use-multilingual-cased-v2` (69.21%) ## **Conclusions** You might be sitting there thinking, "Damn Stuart, that sounds like a hell of a lot of work, is it worth it?". That's a fair question. If you're dealing with small data, this is overkill. However, if you are like us, and you're dealing with big data, then my answer is a resounding "Yes!". And there are four reasons for that: --- ### **It's Better** First and foremost, we found two student models that outperform GPT-4 when compared against the teacher ensemble. These two models were🥇 `sentence-t5-large` (80.40%) and 🥈 `e5-large-v2` (80.21%). Remember the ensemble model achieved 90% agreement with human labels so this is meaningful. ### **It's Cheaper** Second, in our previous post we showed that classifying the sentiment of 10 million news stories would cost an eye-popping \$92,000 using GPT-4. With this student model, which beats GPT-4, we can do the same for \$0.00 marginal cost. In fact, using that $92,000 we could probably buy three Nvidia H100's. ### **It's Faster** Third, using GPT-4 it takes about ~0.5 seconds to classify the sentiment of one news story. 10 million news stories would take 57 compute days. Using the best students models it takes ~0.005 seconds to classify the sentiment of one news story. 10 million stories would take less than 1 compute day to complete. ### **It's Reusable** Finally, you shouldn't see this as a once-off exercise. This is the data science equivalent of a [design pattern](https://en.wikipedia.org/wiki/Software_design_pattern). It can be used and reused over and over again to solve problem after problem. Between ourselves and [Alphix Solutions](https://alphix.com/) we are using this data design pattern to solve a variety of problems like - 1. Sentiment Classification (Positive, Neutral, Negative) 2. Tense Classification (Past, Present, Future) 3. News Classification (M&A, Legal, Results, etc.) 4. Brand Safety Classification (Safe, Unsafe) 5. Beat Classification (Opinion, Formal, Investigative) 6. Political Bias Classification (Left, Centrist, Right) 7. Statement Classification (Forward or Backward looking) 8. IAB Website Classification (Over 700 categories!) --- And that's a wrap. The appendix below contains the code and dataset you need in order to replicate these results for yourself 🤓. ## Appendix This whole analysis was performed in just 198 lines of Python code (less if you don't count lists of model names). Here's everything you need: ``` import datetime as dt import json import numpy as np import pandas as pd from sentence_transformers import SentenceTransformer from sklearn.linear_model import LinearRegression from sklearn.model_selection import train_test_split # Store experiment results. experiment_results_dict = {} # Specify the dataset we want to work with. dataset = "small/nosible-news-small" with open(f"{dataset}.json") as f: # Load the raw data (headlines, etc.) news_data: dict = dict(json.load(f)) with open(f"{dataset}-labels.json", "r") as f: # Load the labels assigned by various LLMs. news_labels: dict = dict(json.load(f)) # Loop through each news story in the dataset. for key, labels in news_labels.items(): ensemble_score = sum( [ # Get the ensemble's score. labels["Text-Unicorn"], labels["GPT-3.5-Turbo"], labels["GPT-4-Original"], labels["Sigma"], labels["Text-Bison"], ] ) news_data[key].update( { # Add the teacher to the dataset. "Teacher": ensemble_score, **labels # Add raw LLMs too. } ) sentences = [] for key, document in news_data.items(): # Create a sentence for each of the news stories we want to train on. sentences.append(f"{document['Headline']}. {document['Description']}") for model_ix, model_name in enumerate( [ # Hugging Face sentence-transformer models. This list goes from models with the # fewest parameters (fastest) to the model with the most parameters (slowest). "TaylorAI/bge-micro-v2", "TaylorAI/bge-micro", "sentence-transformers/all-MiniLM-L6-v2", "sentence-transformers/all-MiniLM-L6-v1", "TaylorAI/gte-tiny", "sentence-transformers/all-MiniLM-L12-v2", "sentence-transformers/all-MiniLM-L12-v1", "BAAI/bge-small-en-v1.5", "thenlper/gte-small", "intfloat/e5-small-v2", "intfloat/e5-small-unsupervised", "sentence-transformers/all-distilroberta-v1", "BAAI/bge-base-en-v1.5", "thenlper/gte-base", "intfloat/e5-base-v2", "intfloat/e5-base-unsupervised", "mukaj/fin-mpnet-base", "sentence-transformers/all-mpnet-base-v2", "sentence-transformers/all-mpnet-base-v1", "sentence-transformers/sentence-t5-base", "sentence-transformers/gtr-t5-base", "intfloat/multilingual-e5-small", "sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2", "sentence-transformers/distiluse-base-multilingual-cased-v1", "sentence-transformers/distiluse-base-multilingual-cased-v2", "intfloat/multilingual-e5-base", "sentence-transformers/paraphrase-multilingual-mpnet-base-v2", "BAAI/bge-large-en-v1.5", "thenlper/gte-large", "intfloat/e5-large-v2", "intfloat/e5-large-unsupervised", "sentence-transformers/sentence-t5-large", "sentence-transformers/gtr-t5-large", "sentence-transformers/all-roberta-large-v1", "intfloat/multilingual-e5-large", ] ): # Load the sentence transformer model. encoder = SentenceTransformer( model_name_or_path=model_name, trust_remote_code=True ) # Count the number of parameters that are in the model. parameters = sum(p.numel() for p in encoder.parameters()) # Start the timer for this model. t0 = dt.datetime.utcnow() # Generate the embeddings. embeddings = encoder.encode( sentences=sentences, show_progress_bar=False ) # Get the dimensionality of the model. dimensionality = embeddings.shape[1] # Get the runtime that it took to encode all sentences. runtime = (dt.datetime.utcnow() - t0).total_seconds() # Create input and output DataFrames for distillation. out_df = pd.DataFrame.from_dict(data=news_data, orient="index") in_df = pd.DataFrame(data=embeddings, index=out_df.index) # Split the data into a training and a validation set. x_train, x_test, y_train, y_test = train_test_split( in_df, out_df, test_size=0.25, random_state=42 ) # Get the out-of-sample teacher labels (Ensemble). teacher_llm = y_test["Teacher"].values teacher_llm_classes = np.zeros(len(teacher_llm)) teacher_llm_classes[teacher_llm <= -1] = -1 teacher_llm_classes[teacher_llm >= 1] = 1 # If this is model #1. if model_ix == 0: for prev_model_name in [ 'TextBlob-0.10', 'TextBlob-0.20', 'TextBlob-0.30', 'VADER-0.10', 'VADER-0.20', 'VADER-0.30', 'Flair-0.50', 'Flair-0.75', 'Flair-0.95', 'Sigma', 'FinBERT', 'FinBERT-Tone', 'Text-Bison', 'Text-Unicorn', 'Gemini-Pro', 'GPT-3.5-Turbo', 'GPT-4-Turbo', 'GPT-4-Original', ]: # Get the predictions of the previous model on the test set. y_test_llm_classes = y_test[prev_model_name].values y_test_llm_sames = (y_test_llm_classes == teacher_llm_classes) y_test_llm_matches = np.sum(y_test_llm_sames) test_accuracy = y_test_llm_matches / len(y_test_llm_sames) * 100 print(f"{prev_model_name} : {test_accuracy:.2f}%") # Add the previous models to the results. experiment_results_dict[prev_model_name] = { "Parameters": np.nan, "Runtime": np.nan, "Dimensions": np.nan, "Accuracy": test_accuracy, } # Fit the linear regression on the embeddings. linear_regression = LinearRegression() linear_regression.fit(x_train.values, y_train["Teacher"].values) # Use the linear regression to predict the teacher sentiment scores. y_pred_test = linear_regression.predict(x_test.values) # Turn the sentiment scores into labels {-1, 0, +1}. y_pred_test = y_pred_test.flatten() y_pred_test_classes = np.zeros(len(y_pred_test)) y_pred_test_classes[y_pred_test <= -1] = -1 y_pred_test_classes[y_pred_test >= 1] = 1 # Calculate the out-of-sample accuracy of the student model. y_test_sames = (y_pred_test_classes == teacher_llm_classes) y_pred_matches = np.sum(y_test_sames) test_accuracy = y_pred_matches / len(y_test_sames) * 100 print(f"OLS({model_name.split('/')[-1]}) : {test_accuracy:.2f}%") # Add the student model to the results with the OLS naming convention. experiment_results_dict[f"OLS({model_name.split('/')[-1]})"] = { "Parameters": parameters, "Runtime": runtime, "Dimensions": dimensionality, "Accuracy": test_accuracy, } # Convert the results into a pandas DataFrame and save it to a CSV file. results_df = pd.DataFrame.from_dict(experiment_results_dict, orient="index") results_df.to_csv("Results.csv") ``` The original 2024 archive has been superseded by the maintained [NOSIBLE Financial Sentiment dataset](https://huggingface.co/datasets/NOSIBLE/financial-sentiment), with 100,000 cleaned, deduplicated, and labelled news samples. ## About Nosible If you're an asset manager looking for ways to add artificial intelligence and large language models into your investment process, we would absolutely love to hear from you. This post, whilst impressive, is just a small demonstration of our full capabilities. The real magic happens at scale and in the intersection. For example, using our index we can extract the rolling 3-month sentiment across all forward-looking statements made about each and every US listed company going back 10 years. If that sounds cool, drop me a line. My email is [stuart@nosible.com](mailto:stuart@nosible.com) 👋 For the current product scope and delivery model, download the [NOSIBLE World Product Overview](https://nosible.com/files/NOSIBLE-World-Product-Overview.pdf). [All Research](https://nosible.com/blog) Related Research ![Running sprinter illustration representing efficient financial-sentiment model training](https://nosible.com/blog/illustrations/the-sprinter.png) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [Financial Sentiment](https://nosible.com/blog/tag/financial-sentiment) ### [Matching GPT-5.1 at Financial Sentiment with Active Learning and Qwen3](https://nosible.com/blog/fast-enough-to-matter-productionizing-tiny-transformers-for-signal-extraction) 2025-12-12 27 min read ![High-resolution chart of selected 35B-A3B serving records across a 350-plus experiment campaign, from Qwen 3.5 FP8 on H200 at $0.218400 per million completion tokens in Experiment 006 to Qwen 3.6 source NVFP4 at $0.125959 in Experiment 332](https://nosible.com/images/2026/08/qwen-35b-a3b-serving-campaign-hero.svg) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) ### [What 350+ Experiments Taught Us About Serving Qwen 3.6 35B-A3B](https://nosible.com/blog/what-350-experiments-taught-us-about-serving-qwen-3-6) 2026-08-17 15 min read ![Thirty semantic factors rendered as daily time series on a dark grid, from geopolitical risk to firm-level political risk.](https://nosible.com/images/2026/08/semantic-factors-hero.png) [Risk Indicator](https://nosible.com/blog/tag/risk-indicator) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ### [The Future That Could Have Been: Turning Web Text into Semantic Stock Betas](https://nosible.com/blog/the-future-that-could-have-been-turning-web-text-into-semantic-stock-betas) 2026-08-05 11 min read > Train a simple regression on sentence embeddings to distil an LLM ensemble—and outperform GPT-4 on financial sentiment classification. **URL:** https://nosible.com/blog/ensemble-and-distil --- --- title: "News Sentiment Showdown: Who Checks Vibes Best?" description: "Compare TextBlob, VADER, FinBERT, Gemini, GPT-3.5 and GPT-4 on 10,368 labelled financial news stories, including the dataset and code." url: "https://nosible.com/blog/news-sentiment-showdown-who-checks-vibes-best" --- [Financial Sentiment](https://nosible.com/blog/tag/financial-sentiment) # News Sentiment Showdown: Who Checks Vibes Best? [Stuart Reid](https://www.linkedin.com/in/stuartgordonreid/) 2024-01-28 12 min read Copy as Markdown ![Abstract eye illustration representing financial-news sentiment analysis](https://nosible.com/blog/illustrations/eye.png) One of our goals for 2024 is to extract investment signals from our growing corpus of verified company news. Last week in the post, " [Using Vector Search to See Signals in Company News](https://nosible.com/blog/using-vector-search-to-see-signals-in-company-news) " we focused on news volume. This week we are going to dive into another dimension we care about: Sentiment. To put it simply, sentiment is how positive or negative the view expressed towards the object or subject of a document is. Quant funds around the world have long been interested in how to quantify sentiment, how sentiment changes over time, and whether it correlates to future market performance. Today we will focus on the first question: "How can sentiment be quantified?". In this post we will compare TextBlob, VADER, Flair, SigmaFSA, FinBERT, FinBERT-Tone, PaLM-2 (Bison and Unicorn), Gemini-Pro, GPT-3.5, GPT-4, and GPT-4-Turbo. In the future we will add more models to the mix. To start off, we compare these models against 250 hand labelled news stories. These stories were chosen at random and serve to answer the question "who checks vibes best". Next, we compare all models side-by-side on a small, but representative, subset of 10,368 news stories and summarize our findings. *I'd like to acknowledge the support of Fundamental Group in the making of this post and, in particular,* [*Jeanne Daniel*](https://twitter.com/jeanniedaniel9) *who helped conduct this analysis. I also want to thank our two December holiday interns - Moabi and Deon - whose project on sentiment model distillation we will discuss next week.* ## Table of Contents 1. [Background](https://nosible.com/blog/news-sentiment-showdown-who-checks-vibes-best#background) 2. [Dataset](https://nosible.com/blog/news-sentiment-showdown-who-checks-vibes-best#dataset) 3. [Models](https://nosible.com/blog/news-sentiment-showdown-who-checks-vibes-best#models) 4. [LLM Prompt](https://nosible.com/blog/news-sentiment-showdown-who-checks-vibes-best#llm-prompt) 5. [Results](https://nosible.com/blog/news-sentiment-showdown-who-checks-vibes-best#results) 6. [Discussion](https://nosible.com/blog/news-sentiment-showdown-who-checks-vibes-best#discussion) 7. [Next Steps](https://nosible.com/blog/news-sentiment-showdown-who-checks-vibes-best#next-steps) 8. [Downloads](https://nosible.com/blog/news-sentiment-showdown-who-checks-vibes-best#downloads) ## Background When ChatGPT launched it quickly became clear that a new state of the art had been achieved across many natural language processing (NLP) tasks. It's fair to say that it upended decades of work and rewrote the playbook. I too had to improvise, adapt and overcome. Since then, the name of the game has been "large language models" and you either played or became irrelevant. One of my favorite patterns in this new post-LLM world is prompting LLMs to label documents and then distilling that capability into smaller, faster, and cheaper models with better guarantees. For those of you who don't know, data labelling is when you assign documents to classes based on their contents. Consider news as an example. Here are some of labels we care about: 1. **Sentiment** - is a news story *Positive*, *Neutral*, or *Negative*? 2. **Forward Looking** - is this *backward looking* or *forward looking*? 3. **Category** - is this a story about *M&A* or a *Charitable Donation*? 4. **Political Bias** - is this story *left-leaning* or *right-leaning*? 5. **Journalistic Beat** - is this an *opinion piece* or *investigative*? Etcetera. From what we have seen the most interesting signals stem from the intersection of one or more classifications. For example, using our system it is possible to ask extremely nuanced questions of the data like: > Is the volume of positive M&A news associated with Microsoft increasing or decreasing relative to the volume of negative M&A news? Or, > Is the sentiment of forward-looking news related to the phrase 'analyst coverage' for General Motors better than that of Ford? In order to answer these questions, you not only need a lot of data (because the intersection shrinks quickly); you also need what I like to call rich datasets. What are rich datasets? They are datasets where everything that ought to be labelled is. Put differently, rich datasets are actually useful to your quants, analysts, data scientists, and software engineers. Poor datasets are not. In the past getting labels was time consuming and expensive. Either you paid per label from sketchy services like Mechanical Turk or you dedicated a month or two to doing it yourself. Now there is a better way. We prompt LLMs. For the rest of this post, we will be focused on sentiment classification but please keep in mind that the same approach also works for other NLP classification tasks. ## Dataset Whether you are training or testing machine learning models, it is extremely important that your dataset is representative of what your model will encounter in the real world. Fortunately, our index makes this much easier to do. - To maximize **dataset representation**, we made sure it contains: - News from all years between 2014 and 2023. - News from all 30 Nosible news categories. - News from all 145 industries in Nosible. - And to maximize **dataset quality**, we only accept news that: - Comes from a financial news source. - Is accepted as related to a given company. - Is not a duplicate of another story. - Does not originate from an AI-generated site. - Was covered by 3+ independent publishers. - Has a title between 5 and 25 words. - Has a description between 5 and 50 words. In practical terms this boiled down to running the following SQL query in a nested for loop and then extracting data for the keys that matched. ``` sql_filter=f""" SELECT key, media_coverage FROM engine WHERE date>='{year + 0}-01-01' AND date<='{year + 1}-01-01' AND nosible_category='{category_name}' AND industry='{industry}' AND source_language_rank>=0.50 AND source_variance_rank>=0.50 AND accepted=true AND apex_story=true AND source_is_banned=false AND media_coverage>=3 AND title_num_words>=5 AND title_num_words<=25 AND description_num_words>=10 AND description_num_words<=50 ORDER BY media_coverage DESC LIMIT 10 """ ``` SQL filter used to produce a high-quality representative dataset of our full News Corpus. Using this filter, we carved out three datasets: - `nosible-news-small` [10,368 records] - Contains the biggest story from each overlapping intersection of year, category, and industry. - `nosible-news-medium` [22,422 records] - Contains the biggest 3 stories from each overlapping intersection of year, category, and industry. - `nosible-news-large` [43,324 records] - Contains the biggest 5 stories from each overlapping intersection of year, category, and industry. Here "biggest" can be defined as - *the story that was the most widely reported on or covered by independent publishers at the same time.* We find that this is a very good proxy for how impactful (and real) a news story is. Each dataset contains the following fields: - **Headline** - The headline of the news story. - **Description** - The Lede or SEO description from the page. - **Category** - The category we assigned it to. - **Company** - The company to which the story is linked. - **Sector** - The GICS sector the linked company. - **Industry** - The GICS industry of the linked company. - **Continent** - The continent the company is from. - **Country** - The country the company is from. Accompanying this post, we are releasing the `nosible-small-news` dataset with the sentiment labels assigned by each of the models we evaluated. This data, and the code we used, is available for download at the end of this blog post. ## Models In this evaluation we tried to include a good mixture of old-school lexicon-based sentiment models, pre-LLM sentiment models, and multiple LLMs. What follows is a list of the models we included and a brief description of each. ### **Random [N/A]** Given a sentence the random classifier returns a -1 (Negative), 0 (Neutral), or 1 (Positive) label with equal probability. I included a random model because they are always useful for computing [post-test probabilities](https://en.wikipedia.org/wiki/Pre-_and_post-test_probability) (left to the reader). ### **TextBlob [2013]** [TextBlob](https://textblob.readthedocs.io/en/dev/) is a useful Python library for text processing. It uses a lexicon of words with associated sentiment scores to label the polarity of sentences. We tested TextBlob using 3x thresholds for the polarity score (0.15, 0.30, and 0.45). ### **VADER [2014]** [VADER](https://www.nltk.org/api/nltk.sentiment.vader.html) (Valence Aware Dictionary and sEntiment Reasoner) is a lexicon and rule-based sentiment analysis model. Given a sentence it computes a compound score that measures sentiment. We tested it with 3x thresholds as well. ### **Flair [2019]** [Flair](https://github.com/flairNLP/flair) is another useful Python library for various NLP tasks. Their pretrained sentiment analysis model uses distilBERT embeddings. Unfortunately, it only outputs two labels - Positive or Negative - so we needed to rejig it. To get Flair to output three classes we looked at the probability assigned to the best label and only accepted that label if the probability crossed a threshold. We tested 3x values for this threshold - 70%, 80%, and 90% confidence. ### **FinBERT [2019]** [FinBERT](https://arxiv.org/abs/1908.10063) is a well-known BERT-based model fine-tuned on financial documents. It was trained to perform three class financial sentiment classification. We used the [ProsusAI FinBERT model card on Hugging Face](https://huggingface.co/ProsusAI/finbert). ### **SigmaFSA [2022]** We couldn't find much on this model, but it looks like a fine-tuned version of Financial BERT (not FinBERT) trained to perform three class financial sentiment classification. We used the [Sigma financial-sentiment-analysis model card on Hugging Face](https://huggingface.co/Sigma/financial-sentiment-analysis). ### **FinBERT Tone [2023]** In 2023 a few finetuned variants of FinBERT were released including ones for ESG, forward looking statement detection, and sentiment. The sentiment one is called FinBERT Tone. Again, we used the Hugging Face [implementation](https://huggingface.co/yiyanghkust/finbert-tone). ### **GPT 3.5 Turbo [2022]** This is the "turbo" version of the original [GPT 3.5 model](https://en.wikipedia.org/wiki/GPT-3) that powered ChatGPT. All of the LLM-based classifiers were given the same prompt. We accessed this model using LangChain. More details on the prompt are given below. ### **Text-Bison [2023]** Text-Bison is one of the [PaLM 2](https://ai.google/discover/palm2) models that Google launched in 2023. This model is optimized for natural language problems and can be accessed through Google Cloud Platform. We kept all the hyperparameters to their defaults. ### **Text-Unicorn [2023]** Text-Unicorn is the largest PaLM 2 model. This model is bigger, more capable, and more expensive than Bison. Again, it can be accessed through Google Cloud Platform. We kept all the hyperparameters to their defaults. ### **GPT 4 [2023]** GPT-4 is the successor model to GPT-3.5. I was made available by OpenAI last year and has demonstrated far superior capabilities across a wide variety of natural language tasks. We accessed GPT-4 through the OpenAI API. ### **GPT 4 Turbo [2023]** GPT-4-Turbo is a more efficient and more cost-effective version of GPT-4. I was made available by OpenAI at the end of last year. We accessed GPT-4-Turbo through the API. Specifically, we used the `gpt-4-1106-preview` model. ### **Gemini-Pro [2023]** Gemini is another suite of large language models released by Google. In this suite there is Gemini-nano, Gemini-pro, and Gemini-ultra. We only have access to Gemini-pro through Google Cloud so we could only include that one. ### **Future Steps** The blog was getting too big, so we stopped at that. In the future we want to follow up with an evaluation of the new Phi-2 LLM from Microsoft as well as Open Source LLMs including various Llama-2, Falcon, and Mistral variants. If you would like to help out, you can reach me at [stuart@nosible.com](mailto:stuart@nosible.com). *P.S. If you're reading this and you work at Google or X - I really want to benchmark Gemini-nano, Gemini-ultra, and Grok on this task. 🧐* ## LLM Prompt When it comes to LLMs, good prompting is crucial. To maximize instruction following ability and accuracy we leveraged few-shot prompting. Zero-shot prompting involves asking an LLM to perform a task unseen. In other words, the LLM is not given any examples of what success or failure looks like. Few-shot prompting, on the other hand, involves asking an LLM to study a few handpicked examples and then, based on what it learns, perform a task. Okay, with that said, here's the few-shot prompt we used for this post. This prompt was used for all of the evaluated LLMs with zero modifications. ``` You are a sentiment classification AI. 1. You are only able to reply with POS, NEU, or NEG. 2. When you are unsure of the sentiment, you MUST reply with NEU. Here are some examples of correct replies: Story: Unexpected demand for the new XYZ product is likely to boost earnings. Reply: POS Story: Sales of the new XYZ product are inline with projections. Reply: NEU Story: Company releases profit warning after the sales of XYZ disappoint. Reply: NEG Story: Following better than expected job numbers, the stock market rallied. Reply: POS Story: The stock market ended flat after job numbers came in as expected. Reply: NEU Story: A large spike in unemployment numbers sent the stock market into panic. Reply: NEG Story: XYZ stock soared after the FDA approved its new cancer treatment. Reply: POS Story: XYZ will announce results of its cancer treatment on the 15th of July. Reply: NEU Story: Following poor results, the FDA shuts down trials of XYZ cancer treatment. Reply: NEG Okay, now please classify the following news story: Story: {story} Reply: ``` ## Results ### **Agreement Matrix - Hand Labelled News** The matrix below shows the percentage of times that each model agreed with every other model across all three classes. The rows and columns have been ordered by how much the model agreed with the hand-labeled human results. ![Agreement matrix comparing sentiment-model labels with human financial-news labels](https://nosible.com/images/2024/01/nosible-human-label-comparison-matrix.png) Agreement matrix comparing sentiment-model labels with human financial-news labels Model agreement matrix for the hand-labeled news. ### **Agreement Matrix - Small News Dataset** The matrix below shows the percentage of times that each model agreed with every other model across all three classes. Here the rows and columns have been ordered by how much they agreed with the best model above - Text-Bison. ![Agreement matrix comparing smaller sentiment models with the best-performing Text-Bison model](https://nosible.com/images/2024/01/nosible-small-label-comparison-matrix.png) Agreement matrix comparing smaller sentiment models with the best-performing Text-Bison model Model agreement matrix for the small news dataset. ### **Cloud LLMs - Small News Dataset Costs** The bar chart below shows the average cost incurred per news story labeled as well as the estimated total cost incurred if we used that model to label 10 million news stories (our corpus contains more documents than this already btw 🫢). ![Bar chart comparing per-story and ten-million-story labeling costs across cloud language models](https://nosible.com/images/2024/01/nosible-small-cloud-llm-costs-1.png) Bar chart comparing per-story and ten-million-story labeling costs across cloud language models ## Discussion Firstly, sentiment classification is not obvious. This is especially true for financial news which often involves multiple entities engaged in complex ways. We went back and forth between ourselves on more than one of the 250 stories we labeled. That said, benchmarking against human labels is the best that we can do. With that disclaimer out of the way, here are some interesting observations: ### **Cloud-Based LLMs win, but at what cost?** - First place went to the Unicorn PaLM-2 LLM from Google. It assigned the same sentiment as we did an incredible 84% of the time! - GPT-4 came second. It agreed with us 74% of the time. This is actually 8% more often than GPT-4-Turbo agreed with our sentiment labels. - GPT-4 and GPT-4-Turbo were (perhaps not too surprisingly) different, and GPT-4-Turbo performed worse than the OG version of GPT-4. - Gemini-Pro and Bison have the same pricing model. If we were choosing between the two for this task, we would choose Gemini-Pro. - Given that Gemini-Pro beat Bison, is it possible that Gemini-Ultra would beat Unicorn? This is speculative, but I'd happily wager on it. - GPT-4-Turbo was 2.98x cheaper than GPT-4. The average cost per story using GPT-4 was \$0.00983 versus\$0.00329 using GPT-4-Turbo. - Whilst their performance is exceptional, we believe that cloud based LLMs are far too expensive at scale for the majority of organizations. ### **FinBERT held up reasonably well.** - If ~69% accuracy is okay, then you are better off using FinBERT than GPT-3.5. It is open source and agrees with Unicorn the same amount. - FinBERT-Tone appears to be a regression rather than an improvement. It only agreed with Unicorn 53% of the time vs. FinBERT's 69%. ### **At 10 years old VADER still kicks!** - VADER with a threshold of 0.10 agreed with Unicorn 56% of the time. This is far better than TextBlob which was very nearly random. - At 56% accurate VADER is not good, but it is 339x faster than FinBERT and may still have a place in real-time trading environments. ### **Flair, SigmaFSA, and FinBERT-Tone** - Despite being far more computationally expensive than VADER, Flair, SigmaFSA, and FinBERT-Tone did not outperform it. Yikes. - Our guess is that these models are probably overfitted to their training datasets or to their domains and are simply not generalizing. ## Next Steps In our next blog post we will show how the performance of massive cloud-based LLMs like GPT-4 and PaLM-2 Unicorn can be distilled into smaller, faster, and cheaper models with better functional and non-functional guarantees. Further down the line we would also like to follow this blog post up with a side-by-side comparison of various open-source large language models. If that sounds interesting to you, we are looking for collaborators for that project. And that's all folks. We are building in public this year so if you liked this content and would like to see more of it, follow us on [twitter](https://twitter.com/nosibleai). If you want to talk, I'm at [stuart@nosible.com](mailto:stuart@nosible.com). ## Downloads Hello Subscriber, 👋 The original 2024 export has been superseded by the maintained [NOSIBLE Financial Sentiment dataset](https://huggingface.co/datasets/NOSIBLE/financial-sentiment), which contains 100,000 cleaned, deduplicated, and labelled news samples. [All Research](https://nosible.com/blog) Related Research ![Running sprinter illustration representing efficient financial-sentiment model training](https://nosible.com/blog/illustrations/the-sprinter.png) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [Financial Sentiment](https://nosible.com/blog/tag/financial-sentiment) ### [Matching GPT-5.1 at Financial Sentiment with Active Learning and Qwen3](https://nosible.com/blog/fast-enough-to-matter-productionizing-tiny-transformers-for-signal-extraction) 2025-12-12 27 min read ![3D cube illustration representing LLM ensemble distillation](https://nosible.com/blog/illustrations/cube.png) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) [Financial Sentiment](https://nosible.com/blog/tag/financial-sentiment) ### [A Pattern for Scaling the Value Proposition of LLMs: Ensemble and Distil 🚀](https://nosible.com/blog/ensemble-and-distil) 2024-02-06 12 min read ![High-resolution chart of selected 35B-A3B serving records across a 350-plus experiment campaign, from Qwen 3.5 FP8 on H200 at $0.218400 per million completion tokens in Experiment 006 to Qwen 3.6 source NVFP4 at $0.125959 in Experiment 332](https://nosible.com/images/2026/08/qwen-35b-a3b-serving-campaign-hero.svg) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) ### [What 350+ Experiments Taught Us About Serving Qwen 3.6 35B-A3B](https://nosible.com/blog/what-350-experiments-taught-us-about-serving-qwen-3-6) 2026-08-17 15 min read > Compare TextBlob, VADER, FinBERT, Gemini, GPT-3.5 and GPT-4 on 10,368 labelled financial news stories, including the dataset and code. **URL:** https://nosible.com/blog/news-sentiment-showdown-who-checks-vibes-best --- --- title: "Using Vector Search to See Signals in Company News" description: "Use vector search over company news to extract investment signals from a multi-terabyte, point-in-time corpus of more than 55 million embeddings." url: "https://nosible.com/blog/using-vector-search-to-see-signals-in-company-news" --- [Web Search](https://nosible.com/blog/tag/web-search) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) [Trading Signals](https://nosible.com/blog/tag/trading-signals) # Using Vector Search to See Signals in Company News [Stuart Reid](https://www.linkedin.com/in/stuartgordonreid/) 2024-01-21 21 min read Copy as Markdown ![Magnifying-glass illustration representing vector search across company news](https://nosible.com/blog/illustrations/inspect.png) The more things change, the more they stay the same. Every year we set ourselves an impossible goal. That has not changed. What has is that we will be building towards that goal *publicly*. So, we invite you to [follow us](https://twitter.com/nosibleai) on this journey and learn alongside us as we tackle a problem of epic proportions. What's the goal? Good question! Our goal is to build a system capable of extracting signals from internet-scale datasets comprised of billions of vectors. When fully realized this system would resemble Google Trends except that it would be more open and finetuned for the investment industry. In this blog post we will take a step back and discuss our 2023 goal. That goal was to curate a high-quality news dataset for all companies and extract investment signals from it. That was a steppingstone to where we are now. The impressive results we achieved are what have inspired us to kick it up a notch. ## Table of Contents - [Table of Contents](https://nosible.com/blog/using-vector-search-to-see-signals-in-company-news#table-of-contents) - [Introduction](https://nosible.com/blog/using-vector-search-to-see-signals-in-company-news#introduction) - [Dataset](https://nosible.com/blog/using-vector-search-to-see-signals-in-company-news#dataset) - [**Infrastructure Setup**](https://nosible.com/blog/using-vector-search-to-see-signals-in-company-news#infrastructure-setup) - [**Company Names**](https://nosible.com/blog/using-vector-search-to-see-signals-in-company-news#company-names) - [**Inline Headlines**](https://nosible.com/blog/using-vector-search-to-see-signals-in-company-news#inline-headlines) - [**AI-Generated News**](https://nosible.com/blog/using-vector-search-to-see-signals-in-company-news#ai-generated-news) - [**Deduplicating News**](https://nosible.com/blog/using-vector-search-to-see-signals-in-company-news#deduplicating-news) - [**Semantic Search**](https://nosible.com/blog/using-vector-search-to-see-signals-in-company-news#semantic-search) - [**Locality Sensitive Hashing**](https://nosible.com/blog/using-vector-search-to-see-signals-in-company-news#locality-sensitive-hashing) - [Signals](https://nosible.com/blog/using-vector-search-to-see-signals-in-company-news#signals) - [**The Code**](https://nosible.com/blog/using-vector-search-to-see-signals-in-company-news#the-code) - [Validation Signals](https://nosible.com/blog/using-vector-search-to-see-signals-in-company-news#validation-signals) - [**Market Signals**](https://nosible.com/blog/using-vector-search-to-see-signals-in-company-news#market-signals) - [Applications](https://nosible.com/blog/using-vector-search-to-see-signals-in-company-news#applications) - [Next Steps](https://nosible.com/blog/using-vector-search-to-see-signals-in-company-news#next-steps) ## Introduction Signals are time series that help us make decisions. The quick and dirty way to measure how informative a signal is, is to look at the difference between the performance of your decisions with that signal versus the performance without that signal. In other words, `G(Decisions|Signal) - G(Decisions)`. Where `G` is an objective measure of performance (a.k.a. the hard part). Signals are usually distilled from larger datasets. These datasets can either be structured or unstructured. For example, [NewMark Risk](https://newmarkrisk.com/) extracts signals from options data (structured). [S3 Partners](https://research.s3partners.com/) is another example, they extract short interest signals from treasury flows (structured). [SwaggyStocks](https://swaggystocks.com/), on the other hand, extracts signals from Reddit comments (unstructured). It is fair to think that because structured data is more available and exhaustively mined by the community, alpha is harder to find and harder to keep secret. For this reason, and because it aligned with our 2023 roadmap, we decided to invest in unstructured data - company news. Given the recent advances in large language models I thought the task would be a little "easier". I was wrong. Take it from me, unstructured data is *still* a pain in the *ss. ## Dataset The largest part of our goal for 2023 was to collect and enrich a corpus of verified company news from around the world. We achieved that goal, but as you will see the journey wasn't easy. At the time of writing, our corpus is a multi-terabyte text and embeddings dataset which contains over 55 million embedded snippets, 150+ million sentences, 4+ billion words, and 5+ billion GPT tokens. For those interested, this dataset is available to all users through various features in our [core product](https://nosible.co.uk/insights/companies/xnas/msft) as well as via API for enterprise customers. More challengingly, the corpus is growing. We add 300,000 new snippets a day and are expanding our coverage of historical news to delisted companies. Adding to a dataset this size whilst maintaining quality and keeping the runtime of our API endpoints under 1 second is difficult. But the data itself was much harder. In no particular order, here are a few of the problems we faced: - Resolving news between companies with nearly identical names. - Stripping out inline headings from inside the body of news stories. - Identifying and filtering out the torrent of AI-generated fake news. - De-duplicating identical news stories across publishers and time. - Approximate nearest neighbors at scale with arbitrary filtering. Because I know you're curious, I'll elaborate a little on each 👀. #### **Infrastructure Setup** But first, it is worth mentioning our data infrastructure. Our ETL system runs across 15 [OVH servers](https://www.ovhcloud.com/en/) in different regions. Each worker is running an API (FastAPI + Gunicorn + Nginx) that receives tasks from Google Cloud Tasks. An example of a task might be "go fetch news about Mastercard". There are various pre-defined tasks available for different entity lists. The worker checks to see if it has the latest news file for Mastercard. If not, the latest file is downloaded from our [Wasabi S3](https://wasabi.com/) bucket. Each file is a compressed [orjson](https://github.com/ijl/orjson) file to reduce latency. Once refreshed, it executes its task and uploads the result to Wasabi S3 and updates our data version control system. Fetching news involves two steps. We find new URLs and then we visit each URL and parse the HTML data. Each request is routed through rotating proxies to ensure that we retrieve the page. Particular emphasis is placed on metadata and structured data in the HTML. Pages without SEO data are often rejected. Similar to Google, we index all data on a regular interval. This involves fetching all the latest news files for all entities and making it searchable. I will talk a little bit more about this indexing process when we introduce the concept of locality sensitive hashing and semantic search further down in this article. #### **Company Names** How many listed companies are called Discovery? Too many. We have Discovery (DSY) in South Africa, Discovery Silver (DSV) in Canada, Discover Financial (DFS) in the USA, discoverIE (DSCV) in the UK, Warner Bros Discovery (WBD) in the US, Predictive Discovery (PDI) in Australia, and the list goes on and on. We solved this using a two-stage model. In the first stage we scan the body of each news story for known keywords associated with the overlapping entities. For example, the names of their CEO and board members. This gives us a dataset of stories about each entity that we have high conviction in. For stage one we used the [Aho-Corasick algorithm implementation](https://github.com/G-Research/ahocorasick_rs) from G-Research. In the second stage we fit probability models. For example, if we see a story about "Discovery" on latimes.com we know there is a high probability of it being about Warner Bros Discovery. Similarly, if we see an article that mentions South Africa, we know it has a nearly 100% probability of being about Discovery (DSY). This works most of the time but still makes bloopers. My personal favorite is Hoya Corp (7741) a Japanese eyewear company. As it turns out, professional boxers like Oscar De la *Hoya* often get hit in the eye and suffer eye injuries. The same kind of eye injuries that *Hoya Corp* can help with. See the problem? Yes, probabilities can be spurious. The best solution is simply to ensemble. #### **Inline Headlines** Have you ever noticed how news websites like to show links to other news stories right in the middle of the article you're reading? It's distracting, I know. One early foible was not stripping those sentences out. Not doing this meant that just about every news story ended up being related to Elon Musk and his escapades. ![Example news article with unrelated inline headlines highlighted as noise for vector search](https://nosible.com/images/2024/01/inlined-headlines.png) Example news article with unrelated inline headlines highlighted as noise for vector search If you're not careful apple watches might like be linked to the ECB cutting interest rates. #### **AI-Generated News** The biggest challenge we faced while curating this corpus was AI generated news. At the beginning of 2023 the problem was small. By July we had "crossed the Rubicon" - more AI news was seen per day than real news. In the image below the AI-generated news is indicated by the red bars. Now you see the problem. ![Time series of Amazon and similar-entity news showing AI-generated articles as red bars](https://nosible.com/images/2024/01/2024-01-20---time-series-of-Amazon.com--AMZN--and-similar-entities.png) Time series of Amazon and similar-entity news showing AI-generated articles as red bars This is a breakdown of news about Amazon including everything except AI-generated news. ![Time series of Amazon and similar-entity news after excluding AI-generated articles](https://nosible.com/images/2024/01/2024-01-20---time-series-of-Amazon.com--AMZN--and-similar-entities---With-Noise.png) Time series of Amazon and similar-entity news after excluding AI-generated articles And here we have included all news that we flagged as being AI-generated. There is a lot to hate about AI generated news. They are usually low-quality knock offs. They don't cite original sources. They are full of inaccuracies and meaningless statements. But worst of all, they break everything. They break search engines, news feeds, social media, public trust, and advertising models. Worst of all, they break signals extracted from company news! We have found that removing AI-generated news improve the signals we extract a lot. Identifying AI-generated news at the story level is extremely difficult. In fact, it might not be possible at all. Ultimately what we ended up doing was blacklisting domains that appear to generate repetitive content in high volumes. To be more specific we look at the distribution of the embeddings of snippets "written" by that domain. If you do that, you'll find that AI-generated sites occupy a relatively small, densely packed region of the vector space when compared to legitimate financial websites such as ft.com, wsj.com and reuters.com. Our theory is that because the articles are generated from a *static* prompt plus some *dynamic* seed content the resulting articles are essentially "constrained" to a small manifold in the vector space. When this metric is combined with a metric of site volume, you get a good predictor. More on this in later blogs. Curious to see what some of these sites look like? Here you go: ![Screenshot of a low-quality website used as an example of AI-generated news](https://nosible.com/images/2024/01/the-spy-who-billed-me.png) Screenshot of a low-quality website used as an example of AI-generated news As you can probably tell even the images for the "news stories" are AI generated. Fortunately, we've noticed two reassuring things when it comes to AI-generated "news" websites. First, quite of few of them have already shut down. I suspect this is because they are operating at a loss. Secondly, we have noticed that Google and Bing are slowly de-indexing these sites from search and news 👏. #### **Deduplicating News** Deduplicating documents is critical for RAG-based applications. Why? Because if you don't de-duplicate your documents you will end up passing multiple permutations of the same text to your LLM. In our experience, this is a great way to damage precision and recall. It is also wasteful. Every time we create a new index, we de-duplicate all news and find the "apex" story for each cluster. De-duplication at scale is hard. Using our ultra-optimized LSH index it is possible to search over all 55-million embeddings in <0.20 seconds. Whilst this is extremely fast - about 10x faster than a flat FAISS index - it is not fast enough. Looping from top to bottom and de-duplicating using the top K nearest neighbors would take an unbelievable 11 million seconds. That's 4 months! Fortunately, we found that duplicate stories are naturally clustered by time and by metadata. Instead of deduplicating across all 55-million embeddings we deduplicate shards of the data that correspond to overlapping periods of time and specific metadata fields. That metadata could be companies, sectors, industries, geographies, or categories. Doing this allows us to deduplicate in minutes. #### **Semantic Search** In our case, semantic search involves finding the N most similar news snippets to a user's search term from the set of news snippets that match their given filters. Filters in this case may include dates, times, authors, publishers, regions, countries, sectors, industries, stocks, and much more. In order to keep it user friendly, filters are expressed through familiar SQL syntax. For example, the following filter: ``` SELECT key FROM engine WHERE date>='2023-01-01' AND date<'2024-01-01' AND apex_story=True AND sector='Technology' ``` tells our index to only return de-duplicated news snippets that occur between the first of January 2023 and the first of January 2024 from the technology sector. Of the 55-million embeddings in our corpus, there are 3,692,137 snippets that match this filter. If we wanted to search over the 3,692,137 vectors associated with those matching snippets, we could do it with the following code snippet: ``` results = esi.search( queries=[ "Adobe Figma Merger" ], page_index=0, page_size=10, sql_filter=""" SELECT key FROM engine WHERE date>='2023-01-01' AND date<'2024-01-01' AND apex_story=True AND sector='Technology' """, ) ``` and the search takes 0.28 seconds on my laptop (faster on our server). Where are those 0.28 seconds spent? It is spent on the following steps: 1. 0.11 seconds - Execute the SQL query to find the right locations. 2. 0.05 seconds - Encode the search term using SBERT (CPU only). 3. 0.09 seconds - LSH Search over the 3,692,137 matching vectors. 4. 0.02 seconds - Reconstruct the vectors to calculate cosine similarity. 5. 0.01 seconds - Compute the cosine similarity of the top N x M results. Let's take a look at some snippets we got back from our above search! 1. *"Adobe to terminate \$20 billion Figma buyout because of regulatory pressure. Adobe Inc and design-tools maker Figma said Monday they have agreed to terminate the \$20 billion merger agreement announced 15 months ago..."* 2. *"Adobe Inc and design-tools maker Figma said Monday they have agreed to terminate the \$20 billion merger agreement announced 15 months ago. Adobe to pay Figma \$1 billion deal-termination fee Adobe and Figma mutually agree to terminate their planned merger amid regulatory pressure..."* 3. *"Adobe ADBE, +2.47%, the maker of Photoshop, Illustrator and other software tools, and Figma said they still believe in the merits of a merger, but they have mutually agreed that there is no clear path to receive the necessary regulatory approvals from the European Commission and the U.K. Competition and Markets Authority. Adobe disclosed that it will pay Figma \$1 billion..."* The top three snippets came from the same apex story. Linked to that Apex story there are more than a dozen other stories that cover the same event. If we had left `apex_story=True` out of our filter we would have near duplicate snippets. It's worth mentioning that this is a little contrived. Most of the time SQL filters are extremely targeted and much faster. For example, a SQL filter that isolates news snippets linked to Adobe from the past quarter executes in less time. #### **Locality Sensitive Hashing** Okay but how does this work? Magic! No, not really. It's just stats. The backbone of our index is a technique called locality sensitive hashing (LSH). To put it simply, LSH works by trying to hash similar items to the same hash code(s). When the order of the elements in a set does not matter and, as a consequence, Jaccard Similarity is a good proxy for similarity, classical LSH algorithms like [SimHash](https://en.wikipedia.org/wiki/Simhash), [MinHash](https://en.wikipedia.org/wiki/MinHash), and [SuperMinHash](https://arxiv.org/abs/1706.05698) are useful. A good example of this kind of situation is good old-fashioned [bag of words](https://en.wikipedia.org/wiki/Bag-of-words_model) (BOW) models. Similarly, when the order of the elements in a set matter and, as a consequence, Hamming Similarity is a good proxy for similarity, LSH algorithms designed to preserve angular distances like [Random Hyperplanes LSH](https://www.cs.princeton.edu/courses/archive/spring04/cos598B/bib/CharikarEstim.pdf), Voronoi LSH, [Cross Polytope LSH](https://arxiv.org/pdf/1602.06922v2.pdf), and [Directional Feature Hashing](https://arxiv.org/pdf/1704.04684v1.pdf) work. We are already over 2,000 words into this article so I will defer explaining more advanced LSH algorithms to future blog posts. I will, however, introduce the concept of Random Hyperplanes LSH. In my opinion, this is an extremely elegant and versatile little algorithm that everybody should know about. Having worked in machine learning for a long time I have come to appreciate the power of visual mental models. So let me introduce you to mine... Imagine that you and I are floating in zero gravity surrounded by thousands of little water bubbles. They are swirling around us on different arced trajectories at different speeds. This is analogous to a 6-dimensional vector space where we have location, $x$, $y$, and $z$, plus velocity, $v$, acceleration, $a$, and plane, $θ$. Now let's imagine that I am tasked with guiding you to a small region of bubbles. How would we do that? We could look at every bubble and its neighborhood one at a time. That's brute force search. Or we could play a game of 20 questions. In this game you ask me twenty questions and I answer "yes" or "no". 1. "Are the bubbles rotating clockwise?" 1. Yes 2. "Are the bubbles speeding up?" 1. No. 3. "Are the bubbles in front of me?" 1. No. 4. "Are the bubbles underneath me?" 1. Yes. 5. "Are the bubbles on a plane of 45° or more?" 1. No. And so on and so forth. With five questions you might already have a decent idea of which bubbles I am referring to. They are behind and underneath you that are rotating clockwise on a plane of more than 45° and are slowing down. More importantly you have a very good idea of which bubbles I am NOT referring to. Rejection matters a lot. This is obviously just a mental model for playing around with ideas, but in many ways 20 questions describes how and why Random Hyperplanes LSH works. In Random Hyperplanes LSH we generate a set of random vectors with the same dimensionality as our dataset. These are the hyperplanes, and they are analogous to your 20 questions. Next, we compute the dot product of each vector in our dataset and those hyperplanes and take the signs; + or -, 0 or 1. These bits are analogous to my answers to your 20 questions. These bits ("answers") tell us which vectors lie on which sides of the various hyperplanes. Similar vectors will tend to fall on the same sides of the same hyperplanes. Finally, we can group these bits together to produce integer hashes that encode the location of vectors. Vectors that produce the same integer hashes are analogous to the bubbles that yield the same answers to your 20 questions. For example, "yes yes no no no yes no yes" → "1 1 0 0 0 1 0 1" -> "197". Just as bubbles that yield the same answers are more likely to be similar, vectors that yield the same integer hashes are more likely to be similar. Using this insight, we can avoid computing cosine similarity on floats and instead compute hamming distances on integers. Because this involves no floating-point operations, most integer operations can be cached, and it is trivial to parallelize, this search can be done extremely quickly on CPUs. This past year we took this concept to an extreme and are now able to search all vectors at a throughput of over 300 million vectors searched per second. ## Signals Okay, let's recap. We now have a high-quality corpus of de-duplicated snippets of financial news plus the ability to filter it by date, time, sector, industry, region, country, company, etc. and search within those filters quickly using LSH. So how do we pull out signals? It's quite simple. First, we apply our filter. Second, we search all vectors. And third, turn the results into a visualization. The benefits of using SBERT embeddings and LSH to extract trends are that: 1. We can do it on the fly for any word, phrase, sentence, or paragraph. 2. It does not require documents to use exact language (keywords). 3. With filters we can extract signals from any subset of documents. 4. The signal can be computed across 55 million embeddings in <1s. #### **The Code** Here is the code we used to pull out the verification signals. To extract signals for specific regions, countries, sectors, industries, companies, and categories of news all we need to do is modify the SQL filter that gets applied to the search. ``` def get_signal(index: EquitySnippetIndex, terms: list, start: dt.date, end: dt.date, sql_filter: str, name: str, triplet=False) -> pd.DataFrame: """ This function accepts a Nosible Snippet index, a list of terms to search for, the start date for the signal, the end date for the signal, a SQL filter to apply to documents, the name of the trend, and whether this is a triplet signal. A triplet signal is a special kind of signal that uses a baseline, a positive, and a negative term. :param index: the Nosible index to search over. :param terms: the terms to search for. :param start: the start date for the signal. :param end: the end date for the signal. :param sql_filter: the SQL filter to apply. :param triplet: whether this is a triplet signal. :param name: the name of the signal. :return: a pandas DataFrame with the signal. """ print(f"EXTRACTING '{name}' SIGNAL") # Add the dates to the SQL filter. sql_filter = f""" {sql_filter} AND date>='{start.strftime('%Y-%m-%d')}' AND date<='{end.strftime('%Y-%m-%d')}' """ # Tidy up the SQL filter so that it prints out nicely. sql_filter = " ".join(sql_filter.split()) # Run the query against the Snippets Polars DataFrame. all_locs = index.get_only_use_locs(sql_filter=sql_filter) date2total = {} for loc in all_locs: date = index.loc2monday[loc] if date not in date2total: date2total[date] = 0 date2total[date] += 1 signal_data = { monday.strftime("%Y-%m-%d"): {term: 0 for term in terms} for monday in pd.date_range(start, end, freq='W-MON') } for term in terms: # Generate an embedding of the trend term using S-BERT. vector = index.vectorize_query(query=term).flatten() # Start the timer. t0 = dt.datetime.utcnow() # Get the most semantically similar snippets. all_sims = index.full_search( vector=vector, locs=all_locs ) # Get the runtime spent in the search method. rt = (dt.datetime.utcnow() - t0).total_seconds() # Remove all results with too low sims. lb = index.max_codes * 0.70 valid_ixs = np.where(all_sims >= lb)[0] all_sims = all_sims[valid_ixs] # Compute the distribution and print out what it looks like. p20, p40, p60, p80, = np.percentile(a=all_sims, q=[20, 40, 60, 80]) for loc, sim in zip(valid_ixs, all_sims): date = index.loc2monday[loc] max_score = date2total[date] * 8 if sim >= p80: signal_data[date][term] += 8 / max_score elif sim >= p60: signal_data[date][term] += 4 / max_score elif sim >= p40: signal_data[date][term] += 2 / max_score elif sim >= p20: signal_data[date][term] += 1 / max_score signal_df = pd.DataFrame.from_dict(signal_data, orient="index") if triplet is False: return signal_df else: # The first signal is the baseline signal. baseline = signal_df[terms[0]] baseline_sma = baseline.rolling(window=4).mean() baseline_mu = baseline.rolling(window=26).mean() baseline_sd = baseline.rolling(window=26).std() baseline = (baseline_sma - baseline_mu) / baseline_sd # The second signal is the positive signal. positive = signal_df[terms[1]] positive_sma = positive.rolling(window=4).mean() positive_mu = positive.rolling(window=26).mean() positive_sd = positive.rolling(window=26).std() positive = (positive_sma - positive_mu) / positive_sd # The third signal is the negative signal. negative = signal_df[terms[2]] negative_sma = negative.rolling(window=4).mean() negative_mu = negative.rolling(window=26).mean() negative_sd = negative.rolling(window=26).std() negative = (negative_sma - negative_mu) / negative_sd # Calculate the true signal by looking at the differences. true_signal = pd.DataFrame(columns=[name], index=signal_df.index) true_signal[name] = (positive - baseline) - (negative - baseline) return true_signal ``` #### Validation Signals In order to verify that the trend system is working as expected let's pull out some signals where we think we know what the signal should look like: ![Seasonal news signal for Christmas, Black Friday, and Valentine's Day mentions](https://nosible.com/images/2024/01/validation-signal-holidays-1.png) Seasonal news signal for Christmas, Black Friday, and Valentine's Day mentions Signal of Christmas, Black Friday, and Valentines mentions (mostly by retailers). ![News signal for COVID, SARS-CoV-2, and coronavirus-related terms](https://nosible.com/images/2024/01/validation-signal-covid.png) News signal for COVID, SARS-CoV-2, and coronavirus-related terms Signal of COVID and COVID-related terms, SARS-CoV-2 and Coronavirus. ![News signal for the Russia-Ukraine war, Ukraine invasion, and Russia sanctions](https://nosible.com/images/2024/01/verification-signal-russia-ukraine-war.png) News signal for the Russia-Ukraine war, Ukraine invasion, and Russia sanctions Signal of Ukraine Invasion and related terms, Russia Ukraine War and Russia Sanctions. ![News signal for Bitcoin, Ethereum, blockchain, and cryptocurrency terms](https://nosible.com/images/2024/01/validation-signal-cryptocurrencies.png) News signal for Bitcoin, Ethereum, blockchain, and cryptocurrency terms Signal of Bitcoin and other cryptocurrency related terms, Ethereum and Blockchain. ![News signal for artificial-intelligence terms](https://nosible.com/images/2024/01/validation-signal-artificial-intelligence.png) News signal for artificial-intelligence terms Signal of AI related terms, Deep Learning, Machine Learning, and Artificial Intelligence. Great, it looks like the system is working because the validation signals look like what we would expect them to look like. This is very reassuring. #### **Market Signals** Whilst playing with this data I wondered to myself whether it would be possible to create new indicators that can tell us how companies in the global economy are performing. As it turns out, we can but it was tricky for three reasons. Firstly, off-the-shelf embedding models are not discriminative enough. To show you what I mean, let's consider the following pair of sentences: - The company *beat* analyst estimates. - The company *missed* analyst estimates. To a human investor these two sentences have very different meanings. To us their semantic similarity is negative. However, to off the shelf embedding models trained on general corpora these sentences look almost exactly the same. Secondly, there is a lot of cyclicality in text that relates to company performance. For example, when you look at the trend of snippets that relate to "beat analyst estimates" you can clearly see when we are in earnings seasons or not. And, thirdly, not only has the volume of company results news increased, but the proportion of company results news has also increased. This lends credibility to the view that investors and journalists are increasingly myopic. Here's a chart that shows you what I mean: ![Investment signal chart comparing three related search terms over time](https://nosible.com/images/2024/01/investment-signal-triplet-terms.png) Investment signal chart comparing three related search terms over time Here we can see that the three components of the triplet signal are correlated, have strong seasonality, and are increasing over time in both absolute and proportionate terms. Nevertheless, using some time series analysis we can still extract a signal. The procedure we used to extract our market barometer signal is as follows: 1. Extract three signals using (1) a baseline term, (2) a positive term, and (3) a negative term. Here are the triplets we considered: - Baseline: "Quarterly Earnings" - Positive: "Quarterly Earnings Grew" - Negative: "Quarterly Earnings Fell" - Baseline: "Analyst Estimates" - Positive: "Beats Analyst Estimates" - Negative: "Misses Analyst Estimates" - Baseline: "Earnings Forecast" - Positive: "Increases Earnings Forecast" - Negative: "Decreases Earnings Forecast" - Baseline: "Quarterly Performance" - Positive: "Strong Quarterly Performance" - Negative: "Weak Quarterly Performance" - Baseline: "Earnings Guidance" - Positive: "Raises Earnings Guidance" - Negative: "Lowers Earnings Guidance" 1. Calculate a rolling Z-score for each of the three terms by subtracting the rolling 6-month average and dividing by the rolling 6-month standard deviation. 2. Then calculate the investment signal as (Positive Z-score - Baseline Z-score) - (Negative Z-score - Baseline Z-score) to get a market indicator. Here we can see the results from 2014-01-01 to the time of writing: ![Investment signal chart showing multiple related terms from 2014 onward](https://nosible.com/images/2024/01/investment-signal-various-terms.png) Investment signal chart showing multiple related terms from 2014 onward Constituents of the market indicator extracted from SBERT embeddings with LSH. And here is what the signal looks like when we take the sum across each of the constituents. Taking the sum makes it easier to see the overall trend. ![Combined investment signal formed by summing the constituent term signals](https://nosible.com/images/2024/01/investment-signal-average.png) Combined investment signal formed by summing the constituent term signals We are currently evaluating these signals and will follow up this blog post with the results. In our evaluation we are looking at how consistent these signals are with existing indicators of economic health. We are also looking at whether the signals are leading or lagging and whether filters can improve them. ## Applications We are also working with one of our enterprise customers, [Sentio Capital](https://www.sentio-capital.com/), to build strategies that leverage our news corpus. We are looking at how signals can be incorporated into risk models. We are also looking at sentiment because it would be nice to split signals up using SQL filters such as the one below: ``` # Look at the signal coming from only positive documents. SELECT key FROM engine WHERE sentiment='Positive' # Look at the signal coming from only negative documents. SELECT key FROM engine WHERE sentiment='Negative' ``` And we are working closely with [Fundamental Group](https://fundamentalgroup.com/), an asset management media buying and planning company and partner in Nosible. We are using signals to understand what the media is talking about and where that conversation is happening. This data is useful for semantic advertising. ``` # Look at the signal coming from only Reuters.com. SELECT key FROM engine WHERE source='reuters.com' # Look at the signal coming from only FT.com and WSJ.com. SELECT key FROM engine WHERE source in ('ft.com', 'wsj.com') ``` It also goes without saying that we are planning on adding this capability to our [core product](https://nosible.co.uk/insights/companies/xnas/msft) in a number of ways. One idea is to provide an LLM context about a company and prompt it to think of signals that may affect it. We could then pull those signals into the platform for users to see and download. ## Next Steps First and foremost, we need to add more data. If you look closely at the holidays signal you will see that the signal gets more refined as time goes by. Unlike some systems, what we have gets *better* as we add more data. So that is what we plan to do. Hence our 2024 goal: push this concept to internet-scale datasets. Secondly, we need to add more diverse data. The language used in SEC filings is different to the language used by journalists in news stories intended to have mainstream appeal. In order to extract signals of more nuanced events that affect companies we must ingest formal text from SEC filings and court cases. Thirdly, we need to train better encoders. Our research has shown that even the best off-the-shelf encoders are not nuanced enough for financial documents. Fortunately, we can use our index to bootstrap a high-quality dataset that we can use to finetune nuanced encoders optimized for financial documents. And that's all folks. If you've stuck with us all the way to the very end, kudos. As I mentioned, we are building in public this year so if you liked this content and would like to see more of it, follow us on [twitter](https://twitter.com/nosibleai). If you want to talk, I'm at [stuart@nosible.com](mailto:stuart@nosible.com). [All Research](https://nosible.com/blog) Related Research ![NOSIBLE World knowledge graph showing entity connections over a decade](https://nosible.com/images/2026/07/kg-hero-decade.png) [Web Search](https://nosible.com/blog/tag/web-search) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) [NOSIBLE World](https://nosible.com/blog/tag/nosible-world) ### [Point-in-Time Knowledge Graphs over Named Entities with NOSIBLE World](https://nosible.com/blog/point-in-time-knowledge-graphs-over-named-entities) 2026-07-16 8 min read ![Abstract vortex illustration representing self-organizing web-scale search facets](https://nosible.com/blog/illustrations/vortex.png) [Web Search](https://nosible.com/blog/tag/web-search) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) ### [Can Faceted Search at Web-Scale Self Organize?](https://nosible.com/blog/can-faceted-search-at-web-scale-self-organize) 2025-10-16 7 min read ![Cybernaut-1 illustration representing agentic search with Monte Carlo Tree Search](https://nosible.com/blog/illustrations/cyber.png) [Web Search](https://nosible.com/blog/tag/web-search) [Information Retrieval](https://nosible.com/blog/tag/information-retrieval) [Artificial Intelligence](https://nosible.com/blog/tag/artificial-intelligence) [Machine Learning](https://nosible.com/blog/tag/machine-learning) ### [Introducing Cybernaut-1: Agentic Search using MCTS](https://nosible.com/blog/introducing-cybernaut-1-agentic-search-with-mcts) 2025-08-26 2 min read > Use vector search over company news to extract investment signals from a multi-terabyte, point-in-time corpus of more than 55 million embeddings. **URL:** https://nosible.com/blog/using-vector-search-to-see-signals-in-company-news --- --- title: "Event ontology index" description: "Explore 16 event ontologies, 4,000+ categories, 48 example usages and downloadable World V1.2 evidence built for investment and risk research." url: "https://nosible.com/ontologies" --- [Home](https://nosible.com/) Ontologies NOSIBLE Research / Reference By NOSIBLE Research Updated 2026-07-18 # Event ontology index Sixteen ontologies turn event reporting into distinct, testable research variables. Raw text does not tell a model whether an event is a threat, sanction, cyber technique, disaster, industry exposure or scheduled information release. These ontologies preserve those distinctions as stable categories that can be filtered, compared and tested through time. Each guide defines the classification boundary, exposes exact World V1.2 counts and uses three example usages to show what the ontology reveals. Start with the research question below, then inspect every category and download the underlying evidence. Ontologies 16 Categories 4,262 Example usages 48 World events 15,311,040 Need the full event payload? [View the complete World event schema.](https://nosible.com/data-dictionaries#world) Choose a research variable ## Open the ontology that matches the question you need to answer [55 categories Flat 100.0% coverage IPTC News Genre Separate what happened from how the information was packaged before it enters an event study. See policy expectations resolve into results, obituaries isolate recycled history, and iPhone reporting move beyond previews. Explore ontology →](https://nosible.com/ontologies/iptc-genre) [5 categories Flat 100.0% coverage Media Frames Measure whether reporting emphasizes costs, blame, conflict, people or morality instead of compressing every narrative into sentiment. See Wirecard move from dispute to accountability, layoffs add blame, and breaches open a liability clock. Explore ontology →](https://nosible.com/ontologies/media-frames) [7 categories Flat 15.2% coverage Ekman-7 Emotion Ontology Preserve threat, blame, loss, rejection and surprise that a positive-negative sentiment score removes. See FTX move from fear to blame, election emotion mean-revert, and COVID reporting pass through distinct crisis states. Explore ontology →](https://nosible.com/ontologies/ekman-emotion) [26 categories Flat 6.6% coverage Schema.org Event Types Turn scheduled public events into comparable information clocks without creating one-off keyword rules for every company. See recurring company conferences, Prime Day openings and entertainment releases form measurable attention windows. Explore ontology →](https://nosible.com/ontologies/schema-org-events) [718 categories 3 levels 100.0% coverage IPTC Media Topics Track when one named event becomes a different economic, political or social story through time. See COVID escape health coverage, Suez persist as logistics risk, and GameStop become a policy story. Explore ontology →](https://nosible.com/ontologies/iptc-media-topics) [709 categories 4 levels 100.0% coverage IAB Content Ontology Convert broad web and media subject matter into stable cohorts that reveal where attention is moving. See regional-bank coverage turn into crisis, strikes disrupt production subjects, and insurance enter hurricane reporting after landfall. Explore ontology →](https://nosible.com/ontologies/iab-content-taxonomy) [237 categories 4 levels 23.0% coverage GICS Industry Classification Measure which industries gain or lose event attention as an economic theme develops through time. See AI concentrate in fewer industries, chip shortages expose automakers, and property stress reach regional banks. Explore ontology →](https://nosible.com/ontologies/gics) [300 categories 2 levels 7.8% coverage ICD-11 Chapters and Blocks Separate health reporting into clinically meaningful cohorts without relying on unstable disease-name keyword lists. See H5N1 gain market context, GLP-1 expand beyond metabolic treatment, and lockdown reporting shift toward clinical disorders. Explore ontology →](https://nosible.com/ontologies/icd-11) [336 categories Flat 8.3% coverage SportsML Categories Measure emerging sports, competitions and recurring attention windows with stable categories instead of brittle keywords. See padel attention nearly triple, Olympic breakdancing create a new cycle, and US Formula One add recurring windows. Explore ontology →](https://nosible.com/ontologies/sportsml) [416 categories 3 levels 29.8% coverage NOSIBLE Event Ontology Follow one company story through changing event mechanisms instead of treating every mention as the same signal. See Credit Suisse move toward resolution, Microsoft-Activision enter integration, and CrowdStrike progress from outage to consequences. Explore ontology →](https://nosible.com/ontologies/nosible-events) [43 categories 3 levels 9.1% coverage EM-DAT Disaster Classification Separate physical hazard regimes and climate-sensitive reporting before testing exposure, pricing or financing effects. See extreme heat track NASA temperatures, climate-sensitive attention rise after normalization, and hazards produce different narrative risks. Explore ontology →](https://nosible.com/ontologies/em-dat) [88 categories 2 levels 11.3% coverage PLOVER Political Events Distinguish political actions that carry different escalation, implementation and market-risk implications. See sanctions displace threats after invasion, UAW actions separate threats from strikes, and concessions follow debt-ceiling agreement. Explore ontology →](https://nosible.com/ontologies/plover) [866 categories 3 levels 9.7% coverage MITRE ATT&CK Enterprise Separate adversary behavior and operational impact so cyber incidents are not reduced to undifferentiated breach counts. See ransomware persistence change, MOVEit move into exfiltration, and three incidents expose different financial-risk channels. Explore ontology →](https://nosible.com/ontologies/mitre-attack) [231 categories 3 levels 100.0% coverage World Geography Measure where an event begins, where attention spreads and which markets inherit the exposure. See Red Sea risk expand across regions, March banking stress cross the Atlantic, and COVID reporting migrate globally. Explore ontology →](https://nosible.com/ontologies/geography) [39 categories 2 levels 10.5% coverage Asset Class Ontology Track which market channels dominate an event narrative as macro regimes and shocks transmit across assets. See inflation synchronize four markets, banking stress reach funding channels, and tariffs split commodity exposure from equity and currency repricing. Explore ontology →](https://nosible.com/ontologies/asset-classes) [186 categories 2 levels 28.1% coverage UN Sustainable Development Goals Separate broad sustainability ambition from specific targets, financing mechanisms and implementation activity. See COP29 shift toward funding, US policy move into clean-energy delivery, and Europe's gas crisis reach demand reduction. Explore ontology →](https://nosible.com/ontologies/sustainable-development-goals) > Explore 16 event ontologies, 4,000+ categories, 48 example usages and downloadable World V1.2 evidence built for investment and risk research. **URL:** https://nosible.com/ontologies --- --- title: "Asset Class Ontology" description: "Explore 39 asset-class categories and example usages tracing inflation, banking contagion and tariff repricing across investable market channels." url: "https://nosible.com/ontologies/asset-classes" --- [Home](https://nosible.com/) [Ontologies](https://nosible.com/ontologies) Asset Classes Ontology field guide By NOSIBLE Research Updated 2026-07-18 # Asset Class Ontology Asset Classes separates market narratives into investable exposure channels that can be compared through time. A market shock rarely stays inside one asset class. Equity losses can trigger margin calls. Investors sell bonds to raise cash. Dollar funding tightens. Commodity prices transmit separate growth and supply shocks. A single finance label hides that sequence and its changing transmission channels. World assigns one main asset class and one investable sub-class to relevant events. Researchers can test whether class composition adds information beyond total coverage and contemporaneous returns. This guide provides every definition, World base rates, three normalized market examples and downloadable observations. [[ 1 ]](https://www.cfainstitute.org/insights/professional-learning/refresher-readings/2026/alternative-investment-features-methods-and-structures) Categories 39 categories Structure 2 levels World events labelled 10.5% Stable codes Since v2 Release data CC0 On this page [01 Foundations](https://nosible.com/ontologies/asset-classes#foundations) [02 Categories](https://nosible.com/ontologies/asset-classes#vocabulary) [03 Example usages](https://nosible.com/ontologies/asset-classes#trends) [04 Downloads and references](https://nosible.com/ontologies/asset-classes#downloads) Foundations ## Asset Classes identifies market exposure without predicting returns, risk or performance The ontology separates equities, fixed income, commodities, currencies, real assets, alternatives, digital assets, and cash and money markets. Thirty-one sub-classes provide investable detail. Funds and derivatives inherit their underlying exposure instead of forming separate classes, preserving the underlying economic risk. [[ 1 ]](https://www.cfainstitute.org/insights/professional-learning/refresher-readings/2026/alternative-investment-features-methods-and-structures) World classifies the asset discussed in an event record. It does not measure holdings, flows, prices or causal exposure. One selected path also compresses multi-asset events. Retain unclassified observations and join independent market data before interpreting a composition change as repricing or allocation. Categories ## Eight asset classes organise 31 sub-classes across investable market exposure types The ontology contains 39 categories. Eight main classes sit above 31 sub-classes. World assigns it to 1,602,795 canonical events, or 10.5% of World V1.2. The explorer exposes every definition, path and corpus count on this page. [[ 1 ]](https://www.cfainstitute.org/insights/professional-learning/refresher-readings/2026/alternative-investment-features-methods-and-structures) Search the ontology 8 / 39 shown World V1.2 coverage Coverage **10.5%** Labelled events **1,602,795** World events **15,311,040** Ontology index 1,602,795 of 15,311,040 World events carry Asset Class Ontology labels. Select a category to inspect the evidence Hierarchy main_class / sub_class › Equities Corporate earnings, duration and risk appetite · 2 children › Fixed Income Rates, duration, credit and market functioning · 5 children › Commodities Physical scarcity, demand and inflation transmission · 5 children › Currencies & FX Relative policy, funding and external balances · 3 children › Real Assets Property, infrastructure and physical value · 3 children › Alternative Investments Private-market and non-traditional exposures · 7 children › Digital Assets Crypto market structure and attention cycles · 3 children › Cash & Money Markets Funding, liquidity and short-term rates · 3 children Equities ### Equities Copy link ↗ #### Definition Ownership stakes in publicly traded companies, including ordinary shares and equity-linked securities. Research use Corporate earnings, duration and risk appetite Code Equities Events with label 253,587 Share of labelled 15.8% Share of World 1.7% **Events with label** is the selected count. **Label share** divides it by 1,602,795 assigned events; **World share** divides it by all 15,311,040 events. #### Representative World V1.2 events One strong classified example per available year, with up to ten years shown. 10 examples 1. 2015-01-08Standard Chartered Cuts 4,000 Jobs, Exits Global Equities BusinessCoverage 95 2. 2016-12-07Asian Shares Rise as Investors Await ECB Response to Italian VoteCoverage 47 3. 2017-03-06Deutsche Bank Shares Plunge, Dragging European Equities LowerCoverage 43 4. 2019-07-25European Equities Fall as Draghi Signals ECB Easing Amid Recession FearsCoverage 28 5. 2020-05-04Asian Equities Plunge as Trump Revives U.S.-China Trade War FearsCoverage 148 6. 2021-10-26Global Equities Rise on Strong Corporate Earnings ReportsCoverage 103 7. 2022-11-22SocGen and Alliance Bernstein Launch Global Equities Joint VentureCoverage 56 8. 2024-03-28Nifty Gains 28 Percent as Indian Equities Close FY24 BullishlyCoverage 62 9. 2025-07-18KuCoin Launches xStocks for Global Tokenized Equities Trading AccessCoverage 149 10. 2026-06-12Binance Launches bStocks Tokenized U.S. Equities With 24/7 TradingCoverage 198 Browse categories to compare definitions, World V1.2 statistics and representative classified events. Example usage 1 of 3 ### Example Usage: Inflation becomes a cross-asset regime as equity, bond, currency and commodity attention roughly doubles Inflation reporting became a synchronized cross-asset regime during the 2022 tightening cycle. Normalized attention to equities rose 116% from its 2021 average, fixed income 92%, currencies 113% and commodities 158%. The Federal Reserve began raising rates in March and delivered its first 75-basis-point increase in June. [[ 2 ]](https://www.federalreserve.gov/newsevents/pressreleases/monetary20220316a.htm) The ontology shows whether a macro narrative remains confined to one market or spreads across several transmission channels. Researchers can use the four normalized series to define inflation-regime windows, test cross-asset correlations, select hedges and compare whether returns, volatility or flows respond when attention becomes synchronized rather than merely elevated. [[ 1 ]](https://www.cfainstitute.org/insights/professional-learning/refresher-readings/2026/alternative-investment-features-methods-and-structures) Equities +116% Fixed income +92% Currencies +113% Commodities +158% **Inflation attention spreads across four investable market channels** Equities Fixed Income Currencies & FX Commodities Monthly category counts divide by all English World events in the same month. The 2021 and 2022 highlights compare annual averages. The cohort requires inflation or monetary-tightening language and an explicit investable-market reference. The chart measures classified attention, not returns, holdings or causal exposure. Example usage 2 of 3 ### Example Usage: Silicon Valley Bank stress reaches venture capital, stablecoins, repo and corporate bonds within fourteen days Silicon Valley Bank closed on 10 March 2023; Signature Bank followed two days later. During the next fourteen days, normalized cohort attention reached 12,912 events per million. Asset subcategories separated the shock into venture-capital, stablecoin, repo and corporate-bond channels rather than one generic banking-crisis label. [[ 3 ]](https://www.fdic.gov/news/press-releases/2023/pr23016.html) [[ 4 ]](https://www.fdic.gov/news/press-releases/2023/pr23017.html) Venture-capital attention peaked first, followed by stablecoins, repurchase agreements and corporate bonds. The sequence gives researchers testable contagion clocks for private funding, digital-asset liquidity, secured funding and credit. Those clocks can be joined to spreads, flows or volatility without assuming that every market channel repriced simultaneously. [[ 4 ]](https://www.fdic.gov/news/press-releases/2023/pr23017.html) Venture capital peak +2,293 per million Stablecoin peak +931 per million Repo peak +637 per million Corporate bonds peak +814 per million Seven-day normalized asset-subcategory attention · per million Venture capital Stablecoins Repo Corporate bonds The fixed English cohort explicitly names Silicon Valley Bank, Signature Bank, First Republic or Credit Suisse. Each line pools subcategory counts across the current and prior six days, then divides by all English World events in the same window. The chart measures reporting channels, not actual holdings or contagion. Example usage 3 of 3 ### Example Usage: April tariffs double normalized equity and currency attention while commodity coverage rises only 38% President Trump announced reciprocal tariffs on 2 April 2025, then suspended most country-specific rates for 90 days on 9 April. Across matched 28-day windows, normalized equity attention rose from 1,659 to 3,500 events per million. Currency attention rose from 1,352 to 2,879, while commodity attention increased from 3,293 to 4,540. [[ 5 ]](https://www.whitehouse.gov/presidential-actions/2025/04/regulating-imports-with-a-reciprocal-tariff-to-rectify-trade-practices-that-contribute-to-large-and-persistent-annual-united-states-goods-trade-deficits/) [[ 6 ]](https://www.whitehouse.gov/presidential-actions/2025/04/modifying-reciprocal-tariff-rates-to-reflect-trading-partner-retaliation-and-alignment/) The divergence distinguishes the trade shock's repricing channels. Equities and currencies more than doubled, while commodity attention rose 38%. Researchers can identify when tariff reporting shifts from goods exposure into valuation and foreign-exchange risk, then test sector returns, implied volatility, basis or hedging demand around that transition. [[ 5 ]](https://www.whitehouse.gov/presidential-actions/2025/04/regulating-imports-with-a-reciprocal-tariff-to-rectify-trade-practices-that-contribute-to-large-and-persistent-annual-united-states-goods-trade-deficits/) Equities +1,841 per million Currencies and FX +1,527 per million Commodities +1,247 per million Seven-day normalized asset-class attention · per million Equities Currencies and FX Commodities The fixed English cohort requires tariff or trade-war language and a United States or major trading-partner reference. Each line pools category counts across the current and prior six days, then divides by all English World events. Highlights compare 5 March–1 April with 2–29 April 2025. Data and sources ## Download Asset Classes and References ### Downloads Download categories and counts as CSV, or the complete machine-readable release as JSON. [**CSV** Definitions and counts Download ↓](https://nosible.com/data/world-v1.2/ontologies/asset-classes/v2/nodes.csv) [**JSON** Full machine-readable release Download ↓](https://nosible.com/data/world-v1.2/ontologies/asset-classes/v2/statistics.json) Release details and citation files [Manifest](https://nosible.com/data/world-v1.2/ontologies/asset-classes/v2/manifest.json) [Release README](https://nosible.com/data/world-v1.2/ontologies/asset-classes/v2/README.md) [CC0 license and scope](https://nosible.com/data/world-v1.2/ontologies/asset-classes/v2/LICENSE.md) [Citation file](https://nosible.com/data/world-v1.2/ontologies/asset-classes/v2/CITATION.cff) [Changelog](https://nosible.com/data/world-v1.2/ontologies/asset-classes/v2/CHANGELOG.md) Need the complete World event schema? [Open the World data dictionary.](https://nosible.com/data-dictionaries#world) ### References 1. [ 1 ][CFA Institute. (2026). Alternative Investment Features, Methods, and Structures. CFA Program Level I curriculum.](https://www.cfainstitute.org/insights/professional-learning/refresher-readings/2026/alternative-investment-features-methods-and-structures) 2. [ 2 ][Board of Governors of the Federal Reserve System. (2022, March 16). Federal Reserve issues FOMC statement.](https://www.federalreserve.gov/newsevents/pressreleases/monetary20220316a.htm) 3. [ 3 ][Federal Deposit Insurance Corporation. (2023). Silicon Valley Bank closure.](https://www.fdic.gov/news/press-releases/2023/pr23016.html) 4. [ 4 ][FDIC, Federal Reserve, and Treasury. (2023). Joint statement.](https://www.fdic.gov/news/press-releases/2023/pr23017.html) 5. [ 5 ][The White House. (2025). Regulating imports with a reciprocal tariff.](https://www.whitehouse.gov/presidential-actions/2025/04/regulating-imports-with-a-reciprocal-tariff-to-rectify-trade-practices-that-contribute-to-large-and-persistent-annual-united-states-goods-trade-deficits/) 6. [ 6 ][The White House. (2025). Modifying reciprocal tariff rates.](https://www.whitehouse.gov/presidential-actions/2025/04/modifying-reciprocal-tariff-rates-to-reflect-trading-partner-retaliation-and-alignment/) Continue exploring ## Complementary ontologies [237 categories GICS Industry Classification Connect market channels to the sectors and sub-industries carrying the underlying exposure. Explore ontology →](https://nosible.com/ontologies/gics) [416 categories NOSIBLE Event Ontology Identify the corporate event mechanisms behind a changing cross-asset narrative. Explore ontology →](https://nosible.com/ontologies/nosible-events) > Explore 39 asset-class categories and example usages tracing inflation, banking contagion and tariff repricing across investable market channels. **URL:** https://nosible.com/ontologies/asset-classes --- --- title: "Ekman-7 Emotion Ontology" description: "Explore seven reported emotion categories, World V1.2 base rates and example usages that separate threat, blame, loss, rejection and surprise through time." url: "https://nosible.com/ontologies/ekman-emotion" --- [Home](https://nosible.com/) [Ontologies](https://nosible.com/ontologies) Ekman-7 Ontology field guide By NOSIBLE Research Updated 2026-07-18 # Ekman-7 Emotion Ontology Ekman-7 separates seven reported emotional frames that one sentiment score collapses and preserves them for direct comparison through time. Positive–negative sentiment hides information. Fear signals threat. Anger signals blame. Sadness signals loss. Disgust signals rejection. Ekman-7 assigns each event one dominant frame: anger, contempt, disgust, fear, happiness, sadness or surprise. The categories come from Ekman’s basic-emotion research. World applies them to reporting, not individual psychology or private emotional states. [[ 1 ]](https://doi.org/10.1080/02699939208411068) [[ 2 ]](https://doi.org/10.1037/h0030377) World applies Ekman-7 to company, government, conflict, hazard, policy and science events. Researchers can test whether each emotion relates to volatility, capital flows, credit spreads or later events. This requires point-in-time data, a defined investment universe, event window and matched baseline. This guide provides definitions, base rates, historical shifts, evidence and downloadable data needed to evaluate the feature before modelling. Categories 7 categories Structure Flat, no hierarchy World events labelled 15.2% Stable codes Since v1 Release data CC0 On this page [01 Foundations](https://nosible.com/ontologies/ekman-emotion#foundations) [02 Categories](https://nosible.com/ontologies/ekman-emotion#vocabulary) [03 Example usages](https://nosible.com/ontologies/ekman-emotion#trends) [04 Downloads and references](https://nosible.com/ontologies/ekman-emotion#downloads) Foundations ## Ekman-7 distinguishes reported emotional frames without inferring the private emotions people feel Ekman treats emotions as distinct response systems, not points on a single positive–negative scale. Each category combines signals, appraisal, context and likely action. Anger indicates obstruction or offence. Fear indicates threat. Disgust indicates rejection. Sadness indicates loss. Happiness indicates desirable outcomes. Surprise indicates new information. Cross-cultural studies support shared category recognition, but expression and interpretation vary by culture. The ontology is a testable classification system, not a tool for inferring private psychological states. [[ 1 ]](https://doi.org/10.1080/02699939208411068) [[ 2 ]](https://doi.org/10.1037/h0030377) [[ 3 ]](https://doi.org/10.1111/j.1467-9280.1992.tb00253.x) World classifies the dominant emotional frame expressed in event reporting. It does not estimate how an author, source or market participant feels. One event may contain several frames, and repeated coverage can increase counts without creating another economic shock. Researchers should test the seven categories against polarity and richer representations. Results should remain stable across sources, sectors, languages, regions and time. Any incremental performance should be linked to specific emotions and market regimes, not attributed to the ontology itself. [[ 4 ]](https://doi.org/10.3389/fpsyg.2018.01217) Categories ## Ekman-7 assigns seven independent emotion labels through a predictable, scalable event classification Ekman-7 is flat: every event receives one label. Anger and contempt are peers, like fear and sadness. Each label describes the dominant frame in reporting, not anyone’s internal state. The explorer shows each category’s definition, code, World V1.2 count and share of labelled events and the full corpus. Use the signals as hypotheses: fear can isolate threat, anger blame, surprise new information and disgust rejection. Search supports ontologies with thousands of categories. Search the ontology 7 / 7 shown World V1.2 coverage Coverage **15.2%** Labelled events **2,334,084** World events **15,311,040** Ontology index 2,334,084 of 15,311,040 World events carry Ekman-7 Emotion Ontology labels. Select a category to inspect the evidence Categories ekman7_emotion anger Blame, confrontation and redress contempt Dismissal, scorn and asserted inferiority disgust Moral or physical rejection and reputational contamination fear Threat, uncertainty and risk-sensitive event cohorts happiness Approval, achievement and favourable outcomes sadness Loss, misfortune and subdued negative response surprise Novel information and abrupt expectation changes anger ### anger Copy link ↗ #### Definition Reporting framed by anger through blame, confrontation, demands for redress or responses to offence or blocked goals. Research use Blame, confrontation and redress Code anger Events with label 1,049,421 Share of labelled 45.0% Share of World 6.9% **Events with label** is the selected count. **Label share** divides it by 2,334,084 assigned events; **World share** divides it by all 15,311,040 events. #### Representative World V1.2 events One strong classified example per available year, with up to ten years shown. 10 examples 1. 2015-04-27Obama Debuts Anger Translator Keegan-Michael Key at Correspondents DinnerCoverage 46 2. 2016-11-12India Banks Struggle to Swap Banned Notes Amid Public AngerCoverage 73 3. 2017-01-06Key and Peele Return for Final Obama Anger Translator Sketch on Daily ShowCoverage 92 4. 2019-11-08Hong Kong Student Dies in Fall During Protests, Sparking AngerCoverage 254 5. 2020-08-06Beirut Blast Kills Over 100, Sparks Government AngerCoverage 998 6. 2021-03-15Newsom Fights Partisan Recall Amid Pandemic and Unemployment AngerCoverage 60 7. 2022-04-19Xi Jinping Maintains Strict Zero-Covid Policy Amid Public AngerCoverage 117 8. 2024-11-04Spain Deploys 7,500 Troops to Flood Zone Amid Rising Public AngerCoverage 401 9. 2025-06-16Justin Bieber Admits Anger Issues After Blocking Friend Over TextsCoverage 261 10. 2026-03-18Rand Paul Confronts Trump's DHS Pick Mullin Over Anger IssuesCoverage 6,756 Browse categories to compare definitions, World V1.2 statistics and representative classified events. Example usage 1 of 3 ### Example Usage: FTX coverage shifts from fear during collapse to anger during prosecution World V1.2 contains 690 Ekman-labelled events tagged with FTX or Sam Bankman-Fried. During the November–December 2022 collapse, fear accounts for 55.3% of 217 events and anger 31.3%. Across the trial window, anger rises to 71.2% while fear falls to 21.6%. Around sentencing, anger reaches 88.9% and fear 7.4%. [[ 5 ]](https://support.ftx.com/hc/en-us/articles/19464725450260-Derivative-Positions) [[ 6 ]](https://www.justice.gov/archives/opa/pr/attorney-general-merrick-b-garland-statement-guilty-verdict-jury-trial-sam-bankman-fried) [[ 7 ]](https://www.justice.gov/usao-sdny/pr/samuel-bankman-fried-sentenced-25-years-prison) Collapse reporting emphasizes threat, insolvency and contagion. Trial and sentencing emphasize blame, wrongdoing and punishment. A 14-day decay preserves the transition without treating each observation as a new regime. This exposes financial-risk mechanisms that a single negative sentiment score cannot distinguish across reported phases in World. [[ 6 ]](https://www.justice.gov/archives/opa/pr/attorney-general-merrick-b-garland-statement-guilty-verdict-jury-trial-sam-bankman-fried) [[ 7 ]](https://www.justice.gov/usao-sdny/pr/samuel-bankman-fried-sentenced-25-years-prison) Collapse Trial Sentencing Full sequence 690 tagged events · 2019-11-06 to 2026-06-17 Collapse · 217 events Fear 55.3% · Anger 31.3% Trial · 111 events Anger 71.2% · Fear 21.6% Sentencing · 27 events Anger 88.9% · Fear 7.4% **Coverage intensity retains a 14-day memory** events per million World events **Coverage moves from threat to blame** fear share minus anger share The decayed series use a 14-day half-life and a seven-event prior estimated only from observations before October 2022. Later observations cannot change the baseline. The longer memory suppresses isolated daily reversals while preserving the collapse, trial and sentencing transitions. This describes media coverage. It does not test returns, causality, attribution or economic impact on exposed assets. Example usage 2 of 3 ### Example Usage: Election day creates a brief happiness spike before anger quickly returns The 2024 United States presidential election produces a sharp but temporary change in classified emotion. Happiness reaches 23.9% in the seven days beginning after election day, up from 6.3% during the prior week. Anger falls from 52.4% to 40.6%, while surprise rises as the outcome becomes known. [[ 9 ]](https://www.fec.gov/resources/cms-content/documents/2024presgeresults.pdf) The transition reverses within two weeks. Happiness falls below 11% and anger returns above 50%, so the improvement has a short half-life. Emotion categories separate outcome confirmation from renewed political conflict. Researchers can then test policy attention, sector exposure, volatility and event-signal decay against each phase. [[ 10 ]](https://doi.org/10.1111/ajps.12819) Happiness +17.6pp Anger -11.8pp Surprise +5.3pp Seven-day pooled share of Ekman-assigned election events · % Happiness Anger Surprise The fixed English cohort covers 15 October through 25 November 2024 and requires election language plus a presidential candidate or voting term. Each line pools Ekman counts across the current and prior six days, then divides by all Ekman-assigned cohort events in that window. Example usage 3 of 3 ### Example Usage: COVID emotion states reveal distinct crisis phases hidden inside negative sentiment COVID coverage doubles after the World Health Organization declares a pandemic on 11 March 2020. Normalized Ekman-labelled attention rises from 50,374 to 102,045 events per million across matched fourteen-day windows. Disgust remains the largest state, but its share falls from 49.3% to 34.9% as the reporting agenda broadens beyond contamination. [[ 8 ]](https://www.who.int/news-room/speeches/item/who-director-general-s-opening-remarks-at-the-media-briefing-on-covid-19---11-march-2020) Anger rises from 6.4% to 16.6%, while sadness triples from 4.5% to 13.5%. Fear falls despite the wider crisis. Those states map to different mechanisms: contamination, policy conflict and realized loss. Researchers can condition volatility, sector exposure and policy-response tests on crisis phase instead of treating every negative COVID event as equivalent. [[ 1 ]](https://doi.org/10.1080/02699939208411068) Normalized coverage +51,671 per million Disgust -14.4pp Anger +10.2pp Sadness +9.0pp Fourteen-day pooled share of Ekman-assigned COVID events · % Disgust Anger Sadness Fear The fixed English COVID cohort compares 26 February to 10 March 2020 with 11 to 24 March. Each line pools emotion counts across the current and prior thirteen days, then divides by all Ekman-assigned cohort events in that window. The marker is the World Health Organization's pandemic declaration. Data and sources ## Download Ekman-7 and References ### Downloads Download categories and counts as CSV, or the complete machine-readable release as JSON. [**CSV** Definitions and counts Download ↓](https://nosible.com/data/world-v1.2/ontologies/ekman-emotion/v1/nodes.csv) [**JSON** Full machine-readable release Download ↓](https://nosible.com/data/world-v1.2/ontologies/ekman-emotion/v1/statistics.json) Release details and citation files [Manifest](https://nosible.com/data/world-v1.2/ontologies/ekman-emotion/v1/manifest.json) [Release README](https://nosible.com/data/world-v1.2/ontologies/ekman-emotion/v1/README.md) [CC0 license and scope](https://nosible.com/data/world-v1.2/ontologies/ekman-emotion/v1/LICENSE.md) [Citation file](https://nosible.com/data/world-v1.2/ontologies/ekman-emotion/v1/CITATION.cff) [Changelog](https://nosible.com/data/world-v1.2/ontologies/ekman-emotion/v1/CHANGELOG.md) Need the complete World event schema? [Open the World data dictionary.](https://nosible.com/data-dictionaries#world) ### References 1. [ 1 ][Ekman, P. (1992). An argument for basic emotions. Cognition & Emotion, 6(3-4), 169-200.](https://doi.org/10.1080/02699939208411068) 2. [ 2 ][Ekman, P., & Friesen, W. V. (1971). Constants across cultures in the face and emotion. Journal of Personality and Social Psychology, 17(2), 124–129.](https://doi.org/10.1037/h0030377) 3. [ 3 ][Ekman, P. (1992). Facial expressions of emotion: New findings, new questions. Psychological Science, 3(1), 34–38.](https://doi.org/10.1111/j.1467-9280.1992.tb00253.x) 4. [ 4 ][Hutto, D. D., Robertson, I., & Kirchhoff, M. D. (2018). A new, better BET: Rescuing and revising Basic Emotion Theory. Frontiers in Psychology, 9, 1217.](https://doi.org/10.3389/fpsyg.2018.01217) 5. [ 5 ][FTX Trading Ltd. (2022, November 11). Voluntary Chapter 11 petitions filed in the United States Bankruptcy Court for the District of Delaware.](https://support.ftx.com/hc/en-us/articles/19464725450260-Derivative-Positions) 6. [ 6 ][U.S. Department of Justice. (2023, November 2). Statement on the guilty verdict in the trial of Sam Bankman-Fried.](https://www.justice.gov/archives/opa/pr/attorney-general-merrick-b-garland-statement-guilty-verdict-jury-trial-sam-bankman-fried) 7. [ 7 ][U.S. Attorney's Office, Southern District of New York. (2024, March 28). Samuel Bankman-Fried sentenced to 25 years in prison.](https://www.justice.gov/usao-sdny/pr/samuel-bankman-fried-sentenced-25-years-prison) 8. [ 9 ][Federal Election Commission. Official 2024 presidential general election results.](https://www.fec.gov/resources/cms-content/documents/2024presgeresults.pdf) 9. [ 10 ][Funck, A. S. (2024). A meta-analytic assessment of the effects of emotions on political information search and decision-making. American Journal of Political Science.](https://doi.org/10.1111/ajps.12819) 10. [ 8 ][World Health Organization. (2020). Director-General's opening remarks at the media briefing on COVID-19, 11 March 2020.](https://www.who.int/news-room/speeches/item/who-director-general-s-opening-remarks-at-the-media-briefing-on-covid-19---11-march-2020) Continue exploring ## Complementary ontologies [5 categories Media Frames Test whether emotion and narrative emphasis carry separate information about the same event. Explore ontology →](https://nosible.com/ontologies/media-frames) [88 categories PLOVER Political Events Connect reported emotion to concrete political actions such as threats, sanctions and concessions. Explore ontology →](https://nosible.com/ontologies/plover) > Explore seven reported emotion categories, World V1.2 base rates and example usages that separate threat, blame, loss, rejection and surprise through time. **URL:** https://nosible.com/ontologies/ekman-emotion --- --- title: "EM-DAT Disaster Classification" description: "Explore 43 EM-DAT disaster categories and evidence linking extreme-heat reporting, normalized climate-sensitive attention and hazard-specific narratives." url: "https://nosible.com/ontologies/em-dat" --- [Home](https://nosible.com/) [Ontologies](https://nosible.com/ontologies) EM-DAT Ontology field guide By NOSIBLE Research Updated 2026-07-18 # EM-DAT Disaster Classification EM-DAT separates disasters by cause, impact channel and financial exposure. Disaster counts are not interchangeable. Heat, floods, storms, earthquakes and industrial accidents create different loss channels, exposed assets and response windows. EM-DAT preserves those distinctions in four levels. World uses a three-level reduction so researchers can construct hazard-specific, point-in-time panels instead of relying on a generic disaster flag. Each cohort retains its distinct transmission mechanism. [[ 1 ]](https://doc.emdat.be/docs/data-structure-and-content/disaster-classification-system/) The distinction matters for asset pricing and risk. Physical hazards affect property, production, insurance, credit and public balance sheets through different mechanisms. A classified event stream can show when each mechanism enters reporting and which entities or regions are exposed. It cannot replace physical observations, loss estimates or causal attribution from independent evidence. [[ 4 ]](https://www.ngfs.net/en/press-release/ngfs-publishes-latest-long-term-climate-macro-financial-scenarios-climate-risks-assessment) [[ 5 ]](https://doi.org/10.1111/jofi.13219) Categories 43 categories Structure 3 levels World events labelled 9.1% Stable codes Since v1 Release data CC0 On this page [01 Foundations](https://nosible.com/ontologies/em-dat#foundations) [02 Categories](https://nosible.com/ontologies/em-dat#vocabulary) [03 Example usages](https://nosible.com/ontologies/em-dat#trends) [04 Downloads and references](https://nosible.com/ontologies/em-dat#downloads) Foundations ## EM-DAT separates disaster mechanisms but does not measure physical hazard intensity EM-DAT divides disasters into natural and technological groups, then into subgroups, types and subtypes. World stops at type. The hierarchy separates meteorological, hydrological, climatological, geophysical and biological events. This matters because heat, floods, wildfire and earthquakes have different spatial footprints, lead times and financial transmission channels. Those differences determine the appropriate exposure map. [[ 1 ]](https://doc.emdat.be/docs/data-structure-and-content/disaster-classification-system/) [[ 2 ]](https://www.ipcc.ch/report/ar6/wg1/chapter/chapter-11/) World classifies the event described in reporting. It does not determine whether an event meets the loss or mortality thresholds used by the EM-DAT disaster database. It also does not estimate physical intensity, insured loss or climate attribution. Validate every result against independent hazard, exposure and vulnerability data. [[ 1 ]](https://doc.emdat.be/docs/data-structure-and-content/disaster-classification-system/) Categories ## EM-DAT uses three levels to distinguish natural, technological and complex disasters World implements two roots, nine intermediate groups and 32 disaster types across three levels. Each assigned event receives one path from group to type. The explorer exposes definitions, stable codes, full paths and World V1.2 counts. Search and bounded tree expansion keep the interface usable as larger ontologies are added. Search the ontology 2 / 43 shown World V1.2 coverage Coverage **9.1%** Labelled events **1,393,908** World events **15,311,040** Ontology index 1,393,908 of 15,311,040 World events carry EM-DAT Disaster Classification labels. Select a category to inspect the evidence Hierarchy disaster_group / disaster_subgroup / disaster_type › Natural Natural hazards grouped by physical mechanism · 6 children › Technological Human, infrastructure and industrial failure events · 3 children Natural ### Natural Copy link ↗ #### Definition Disasters caused by geophysical, meteorological, hydrological, climatological, biological or extraterrestrial hazards. Research use Natural hazards grouped by physical mechanism Code Natural Events with label 661,379 Share of labelled 47.4% Share of World 4.3% **Events with label** is the selected count. **Label share** divides it by 1,393,908 assigned events; **World share** divides it by all 15,311,040 events. #### Representative World V1.2 events One strong classified example per available year, with up to ten years shown. 10 examples 1. 2015-10-04French Riviera Floods Kill At Least 16, Natural Disaster DeclaredCoverage 95 2. 2016-11-14World Bank: Natural Disasters Cost $520bn and Push 26m into PovertyCoverage 40 3. 2017-01-042016 Natural Disasters Cause Record $175 Billion in Global DamagesCoverage 20 4. 2019-01-22Poll: Natural Disasters Shape American Climate Change ViewsCoverage 29 5. 2020-10-12UN: Climate Change Doubled Natural Disasters Since 2000Coverage 29 6. 2021-09-25Haitian Migrants Flee Natural Disasters Amid Global Focus on Del RioCoverage 27 7. 2022-09-12Kentucky Natural Disasters Surge in Frequency and CostCoverage 190 8. 2024-02-28Natural Disasters and Inflation Drive Rising Home Auto Insurance CostsCoverage 141 9. 2025-01-24US Natural Disaster Losses Hit $218 Billion in 2024Coverage 72 10. 2026-01-24US Winter Storm Cuts Natural Gas Output and Spikes Power PricesCoverage 55 Browse categories to compare definitions, World V1.2 statistics and representative classified events. Example usage 1 of 3 ### Example Usage: Extreme-heat reporting nearly doubles as NASA temperatures rise while cold coverage barely moves World's EM-DAT Extreme temperature category combines heat and cold events at one level. We split 23,139 English assignments into explicit heat, cold or unresolved reporting cohorts, then divide monthly counts by all English World events. Heat rises from 81.0 to 160.9 reports per 100,000 between 2015-2019 and 2023-2025. [[ 1 ]](https://doc.emdat.be/docs/data-structure-and-content/disaster-classification-system/) NASA's global anomaly rises from 0.93°C to 1.21°C across the same periods. Heat coverage increases 98.7%; cold and severe-winter coverage increases 25.5%. The smoothed month-to-month relationship is modest (r = 0.22), so the evidence supports structural alignment, not a claim that global temperature mechanically determines monthly reporting throughout the geographically heterogeneous World reporting corpus. [[ 6 ]](https://data.giss.nasa.gov/gistemp/index.html) [[ 7 ]](https://www.nasa.gov/news-release/temperatures-rising-nasa-confirms-2024-warmest-year-on-record/) Heat reporting +98.7% Cold reporting +25.5% NASA anomaly +0.28°C Monthly relationship r = 0.22 **Heat coverage separates from cold as the measured climate warms** monthly rate per 100,000 English World events · two-month half-life Heat reporting Cold reporting NASA anomaly **Heat reporting rises 99%; cold reporting rises 26%** pooled rate per 100,000 English World events **Heat** 2015-2019 81.0 2023-2025 160.9 **Cold and severe winter** 2015-2019 39.5 2023-2025 49.5 Heat and cold are mutually exclusive analytical cohorts inside World's EM-DAT Extreme temperature assignments, not additional EM-DAT categories. Rates use all English World events as the denominator. NASA GISTEMP is a global temperature measure, so the lines test co-movement in reporting rather than local weather attribution or direct measures of local hazard exposure across entities, assets and locations worldwide. Example usage 2 of 3 ### Example Usage: Climate-sensitive disaster coverage rises 44% after normalization for the growing World corpus World assigns 21,706 events to Drought, Extreme temperature or Flood in 2015-2019 and 89,211 in 2023-2025. Raw growth is not comparable because the World corpus also expanded. After dividing by all events, the basket rises from 901.4 to 1,296.7 assignments per 100,000. That is a 43.9% increase beyond corpus growth. [[ 1 ]](https://doc.emdat.be/docs/data-structure-and-content/disaster-classification-system/) Floods damage property and infrastructure. Extreme temperatures affect labour, power demand and physical assets. Droughts constrain agriculture, water and hydroelectric generation. The chart shows when coverage of these financial exposures rises. Researchers can then join each hazard to locations, issuers, insurers and sovereign balance sheets. [[ 2 ]](https://www.ipcc.ch/report/ar6/wg1/chapter/chapter-11/) [[ 3 ]](https://doi.org/10.3386/w30445) 2015-2019 901.4 per 100k 2023-2025 1296.7 per 100k Normalized increase +43.9% Classified events 151,772 **Normalized climate-sensitive disaster coverage rises 44%** rate per 100,000 World events · two-month half-life **Extreme temperature contributes the largest normalized increase** pooled rate per 100,000 World events **Extreme temperature** 2015-2019 395.8 2023-2025 548.4 **Flood** 2015-2019 296.5 2023-2025 428.2 **Drought** 2015-2019 209.1 2023-2025 320.1 The basket contains World events assigned to EM-DAT Drought, Extreme temperature or Flood. Monthly counts are divided by all World events and smoothed with a two-month half-life. The chart measures classified reporting, not physical disaster incidence, climate attribution or investment performance. Example usage 3 of 3 ### Example Usage: Earthquakes emphasize economic impact while wildfires concentrate responsibility and conflict across reporting EM-DAT categories separate hazards that a generic disaster flag combines. In a fixed first-quarter 2024 English cohort, Economic Consequences frames 78.5% of earthquake events but only 21.0% of wildfire events. Responsibility leads wildfire coverage at 37.4%, while Conflict reaches 38.0% for drought, across four preselected hazard cohorts. [[ 1 ]](https://doc.emdat.be/docs/data-structure-and-content/disaster-classification-system/) [[ 8 ]](https://doi.org/10.1111/j.1460-2466.2000.tb02843.x) Those differences change the research question. Earthquakes create abrupt asset-damage and reconstruction windows. Drought creates persistent water, food and power constraints. Wildfire reporting emphasizes liability as well as physical loss. Hazard-specific classification therefore determines which entities, exposures and event horizons belong in a risk panel before returns, losses, spreads or portfolio outcomes are tested. [[ 4 ]](https://www.ngfs.net/en/press-release/ngfs-publishes-latest-long-term-climate-macro-financial-scenarios-climate-risks-assessment) Earthquake economic frame 78.5% Wildfire responsibility 37.4% Drought conflict 38.0% Wildfire non-economic 79.0% **Each hazard produces a different narrative-risk mix** share of classified English events · Q1 2024 Responsibility Economic Consequences Conflict Human Interest Morality **Wildfire** 377 **Flood** 237 **Earthquake** 247 **Drought** 805 0% 100% · event count at right **Only 21% of wildfire coverage is economically framed, versus 79% for earthquakes** share outside Economic Consequences **Wildfire** 79.0 % **Flood** 45.6 % **Earthquake** 21.5 % **Drought** 51.9 % The fixed cohort contains English World events from January through March 2024 with both an EM-DAT disaster-type assignment and a Media Frames assignment. Bars show within-hazard frame shares. They describe how reporting packages risk; they do not estimate losses, public sentiment or causal market effects. Data and sources ## Download EM-DAT and References ### Downloads Download categories and counts as CSV, or the complete machine-readable release as JSON. [**CSV** Definitions and counts Download ↓](https://nosible.com/data/world-v1.2/ontologies/em-dat/v1/nodes.csv) [**JSON** Full machine-readable release Download ↓](https://nosible.com/data/world-v1.2/ontologies/em-dat/v1/statistics.json) Release details and citation files [Manifest](https://nosible.com/data/world-v1.2/ontologies/em-dat/v1/manifest.json) [Release README](https://nosible.com/data/world-v1.2/ontologies/em-dat/v1/README.md) [CC0 license and scope](https://nosible.com/data/world-v1.2/ontologies/em-dat/v1/LICENSE.md) [Citation file](https://nosible.com/data/world-v1.2/ontologies/em-dat/v1/CITATION.cff) [Changelog](https://nosible.com/data/world-v1.2/ontologies/em-dat/v1/CHANGELOG.md) Need the complete World event schema? [Open the World data dictionary.](https://nosible.com/data-dictionaries#world) ### References 1. [ 1 ][Centre for Research on the Epidemiology of Disasters. EM-DAT disaster classification system, 2023 release.](https://doc.emdat.be/docs/data-structure-and-content/disaster-classification-system/) 2. [ 2 ][Seneviratne, S. I. et al. (2021). Weather and climate extreme events in a changing climate. In IPCC AR6 Working Group I, Chapter 11.](https://www.ipcc.ch/report/ar6/wg1/chapter/chapter-11/) 3. [ 3 ][Acharya, V. V. et al. (2022). Is physical climate risk priced? Evidence from regional variation in exposure to heat stress. NBER Working Paper 30445.](https://doi.org/10.3386/w30445) 4. [ 4 ][Network for Greening the Financial System. (2024). NGFS long-term climate macro-financial scenarios, Phase V.](https://www.ngfs.net/en/press-release/ngfs-publishes-latest-long-term-climate-macro-financial-scenarios-climate-risks-assessment) 5. [ 5 ][Sautner, Z., van Lent, L., Vilkov, G., & Zhang, R. (2023). Firm-level climate change exposure. Journal of Finance, 78(3), 1449-1498.](https://doi.org/10.1111/jofi.13219) 6. [ 6 ][NASA Goddard Institute for Space Studies. GISS Surface Temperature Analysis version 4 (GISTEMP v4).](https://data.giss.nasa.gov/gistemp/index.html) 7. [ 7 ][NASA. (2025). Temperatures Rising: NASA Confirms 2024 Warmest Year on Record.](https://www.nasa.gov/news-release/temperatures-rising-nasa-confirms-2024-warmest-year-on-record/) 8. [ 8 ][Semetko, H. A., & Valkenburg, P. M. (2000). Framing European politics: A content analysis of press and television news. Journal of Communication, 50(2), 93-109.](https://doi.org/10.1111/j.1460-2466.2000.tb02843.x) Continue exploring ## Complementary ontologies [231 categories World Geography Locate disaster exposure by continent, region and country before joining it to assets or entities. Explore ontology →](https://nosible.com/ontologies/geography) [5 categories Media Frames Measure whether each hazard is presented through cost, responsibility, conflict or human impact. Explore ontology →](https://nosible.com/ontologies/media-frames) > Explore 43 EM-DAT disaster categories and evidence linking extreme-heat reporting, normalized climate-sensitive attention and hazard-specific narratives. **URL:** https://nosible.com/ontologies/em-dat --- --- title: "World Geography" description: "Explore 231 geographic categories and example usages tracing Red Sea shipping risk, banking stress and pandemic reporting across regions and countries." url: "https://nosible.com/ontologies/geography" --- [Home](https://nosible.com/) [Ontologies](https://nosible.com/ontologies) Geography Ontology field guide By NOSIBLE Research Updated 2026-07-18 # World Geography World Geography shows where an event occurs, allowing local shocks and cross-border propagation to be measured separately through time. A source location is not an event location. A London publisher can report an attack in Yemen, a factory stoppage in Germany or a freight response in India. Collapsing those places into the publisher's country destroys the transmission path. World Geography assigns each event a continent, region and country using stable geographic labels. The resulting path follows the reported event itself through time. [[ 1 ]](https://www.iso.org/iso-3166-country-codes.html) [[ 2 ]](https://unstats.un.org/unsd/methodology/m49/) Researchers can use the hierarchy to construct local cohorts, measure geographic diffusion and join events to country, port, trade or asset exposures. The field does not identify publisher domicile, issuer revenue geography or causal spillovers. This guide provides the full structure, World base rates, a Red Sea shipping study and data for testing geographic event signals against independently defined exposures. Categories 231 categories Structure 3 levels World events labelled 100.0% Stable codes Since v1 Release data CC0 On this page [01 Foundations](https://nosible.com/ontologies/geography#foundations) [02 Categories](https://nosible.com/ontologies/geography#vocabulary) [03 Example usages](https://nosible.com/ontologies/geography#trends) [04 Downloads and references](https://nosible.com/ontologies/geography#downloads) Foundations ## World Geography classifies event locations without inferring source or issuer exposure The ontology follows a three-level continent, region and country hierarchy grounded in ISO country codes and United Nations regional groupings. Stable country labels support joins across time, while the parent path allows aggregation when a country sample is too sparse. Repeated regional labels on different paths remain explicit in the explorer. [[ 1 ]](https://www.iso.org/iso-3166-country-codes.html) [[ 2 ]](https://unstats.un.org/unsd/methodology/m49/) World assigns the event location. It does not identify publisher location, exchange venue, company domicile or economic exposure. A factory event in Germany can affect suppliers and securities elsewhere. Join each event location to a separate, point-in-time exposure map before testing spillovers. Categories ## World Geography organises 231 categories into continents, regions and countries The hierarchy contains six continent roots, regional groupings and 194 country leaves across three levels. Every World event receives one geographic path. The explorer exposes definitions, stable codes, full paths and World V1.2 counts. Search and bounded expansion keep the full geography accessible without separate country pages. [[ 1 ]](https://www.iso.org/iso-3166-country-codes.html) [[ 2 ]](https://unstats.un.org/unsd/methodology/m49/) Search the ontology 6 / 231 shown World V1.2 coverage Coverage **100.0%** Labelled events **15,311,040** World events **15,311,040** Ontology index 15,311,040 of 15,311,040 World events carry World Geography labels. Select a category to inspect the evidence Hierarchy continent / region / country › Asia Events located in Asia or its regional groups · 9 children › Europe · 8 children › Africa · 8 children › North America Global market, security and shipping-policy responses · 4 children › South America · 5 children › Oceania · 14 children Asia ### Asia Copy link ↗ #### Definition Events geographically associated with the continent Asia. Research use Events located in Asia or its regional groups Code Asia Events with label 5,097,396 Share of labelled 33.3% Share of World 33.3% **Events with label** is the selected count. **Label share** divides it by 15,311,040 assigned events; **World share** divides it by all 15,311,040 events. #### Representative World V1.2 events One strong classified example per available year, with up to ten years shown. 10 examples 1. 2015-10-26Deadly Earthquake Rocks Northern Afghanistan Killing Over 200 Across AsiaCoverage 658 2. 2016-08-25Apple Issues iOS Update After Spyware Targets Activist in West AsiaCoverage 195 3. 2017-11-05Trump kicks off Asia tour in Japan with tough North Korea stanceCoverage 511 4. 2019-12-26Thousands Witness Rare Ring of Fire Solar Eclipse Across AsiaCoverage 344 5. 2020-07-09BCCI President Ganguly Announces Asia Cup T20 Cancellation Due to COVID-19Coverage 379 6. 2021-08-20Harris Asia Trip Gains Urgency After Afghan Taliban TakeoverCoverage 176 7. 2022-06-14Biden Visits Israel and Saudi Arabia to Launch West Asia QuadCoverage 868 8. 2024-08-09Japan PM Cancels Asia Trip After Megaquake Advisory IssuedCoverage 1,608 9. 2025-09-29India Refuses Asia Cup 2025 Trophy After Beating Pakistan in Dubai FinalCoverage 4,437 10. 2026-03-23PM Modi Addresses Lok Sabha on West Asia Crisis and Oil Supply RisksCoverage 1,514 Browse categories to compare definitions, World V1.2 statistics and representative classified events. Example usage 1 of 3 ### Example Usage: Red Sea reporting expands from three regions to six as carriers reroute shipping The cohort contains 1,766 English-language events that mention a Red Sea, Houthi, Bab el-Mandeb or Suez identifier and a maritime-shipping term. Outside-Middle-East event share rises from 27.9% during initial attacks to 46.7% during carrier rerouting. Effective regions rise from 2.95 to 5.56, revealing diffusion that global counts miss. [[ 3 ]](https://www.imo.org/en/mediacentre/pressbriefings/pages/imo-msc-resolution-red-sea.aspx) Carrier rerouting provides the mechanism for geographic diffusion. Weekly changes in operational and ticker-linked financial coverage show a weak six-week relationship (r = 0.17), but the lag-adjusted test is not significant (p = 0.119). The evidence supports broader regional exposure, not a stable propagation delay, trade effect or price response. [[ 4 ]](https://unctad.org/publication/navigating-troubled-waters-impact-global-trade-disruption-shipping-routes-red-sea-black) [[ 5 ]](https://www.ecb.europa.eu/press/projections/html/ecb.projections202403_ecbstaff~f2f2d34d5a.en.html) [[ 6 ]](https://openknowledge.worldbank.org/bitstreams/de1a433c-e9ab-4eb3-ae7c-430c3a463d91/download) Outside Middle East **27.9 % → 46.7 %** Effective regions **2.95 → 5.56** Lag-adjusted test **p = 0.119** **Regional breadth doubles as carriers reroute around the Red Sea** effective regions represented each month Monthly coverage divides the fixed Red Sea shipping cohort by all English World events. Regional measures use shares within that cohort. The lead test compares weekly changes in same-region normalized rates and corrects for selecting the strongest lag. The study establishes reporting diffusion, not trade, inflation, freight-rate or return effects. Those remain separate questions for external economic data and cannot be inferred from coverage alone. Example usage 2 of 3 ### Example Usage: March banking stress cuts North America’s share as Western Europe quadruples March 2023 banking stress begins as a North American event and rapidly acquires a European dimension. Normalized coverage rises from 546 to 917 events per million after 13 March. North America falls from 79.7% to 47.1% of assigned geography, while Western Europe rises from 9.8% to 40.4%. [[ 7 ]](https://www.federalreserve.gov/publications/2023-may-financial-stability-report-funding-risks.htm) The geographic mix distinguishes local failure from cross-border transmission. Silicon Valley Bank and Signature Bank drive the first phase; Credit Suisse dominates the next. Effective regional concentration rises from 1.54 to 2.58 regions. That change separates liquidity stress from a cross-border banking-risk regime before price, funding and policy tests. [[ 8 ]](https://oig.federalreserve.gov/reports/board-material-loss-review-silicon-valley-bank-sep2023.htm) Western Europe +30.6pp North America -32.6pp Effective regions +1.04 Normalized coverage +371 per million March banking stress cuts North America’s share as Western Europe quadruples 8 to 12 March 13 to 31 March 9.8% 40.4% **Western Europe** 79.7% 47.1% **North America** The fixed cohort explicitly names Silicon Valley Bank, Signature Bank, First Republic or Credit Suisse. It contains 123 assigned events before 13 March and 940 afterwards across the matched North American and European stress phases in March. Example usage 3 of 3 ### Example Usage: COVID-19 coverage moves from East Asia into North America and South Asia COVID-19 reporting moves decisively out of East Asia as global transmission accelerates. East Asia falls from 40.9% to 3.7% of assigned geography after 20 February 2020. North America rises from 18.8% to 48.9%, South Asia reaches 18.8%, and normalized cohort coverage increases from 33,683 to 409,295 events per million. [[ 9 ]](https://www.who.int/emergencies/disease-outbreak-news/item/2020-DON233) The result measures a shift in reporting geography, not wider regional diversity. Attention moves toward North America, South Asia and Western Europe as transmission accelerates. Geography labels expose that redistribution and create region-specific event clocks for testing travel restrictions, supply-chain disruption, policy responses and cross-market transmission around the pandemic declaration in early 2020. [[ 10 ]](https://www.who.int/news-room/speeches/item/who-director-general-s-opening-remarks-at-the-media-briefing-on-covid-19---11-march-2020) East Asia -37.2pp North America +30.1pp South Asia +8.1pp Normalized coverage +375,612 per million Weekly share of assigned geography · % East Asia North America South Asia The fixed English cohort requires COVID, COVID-19, coronavirus or SARS-CoV-2. Weekly shares use Geography-assigned cohort events; normalized attention divides by all English World events. The pre-registered boundary is 20 February 2020. Data and sources ## Download Geography and References ### Downloads Download categories and counts as CSV, or the complete machine-readable release as JSON. [**CSV** Definitions and counts Download ↓](https://nosible.com/data/world-v1.2/ontologies/geography/v1/nodes.csv) [**JSON** Full machine-readable release Download ↓](https://nosible.com/data/world-v1.2/ontologies/geography/v1/statistics.json) Release details and citation files [Manifest](https://nosible.com/data/world-v1.2/ontologies/geography/v1/manifest.json) [Release README](https://nosible.com/data/world-v1.2/ontologies/geography/v1/README.md) [CC0 license and scope](https://nosible.com/data/world-v1.2/ontologies/geography/v1/LICENSE.md) [Citation file](https://nosible.com/data/world-v1.2/ontologies/geography/v1/CITATION.cff) [Changelog](https://nosible.com/data/world-v1.2/ontologies/geography/v1/CHANGELOG.md) Need the complete World event schema? [Open the World data dictionary.](https://nosible.com/data-dictionaries#world) ### References 1. [ 1 ][International Organization for Standardization. ISO 3166 country codes.](https://www.iso.org/iso-3166-country-codes.html) 2. [ 2 ][United Nations Statistics Division. Standard country or area codes for statistical use (M49).](https://unstats.un.org/unsd/methodology/m49/) 3. [ 3 ][International Maritime Organization. (2024, May 24). IMO condemns illegal, unjustifiable attacks on ships in Red Sea.](https://www.imo.org/en/mediacentre/pressbriefings/pages/imo-msc-resolution-red-sea.aspx) 4. [ 4 ][UN Trade and Development. (2024). Navigating troubled waters: Impact to global trade of disruption of shipping routes in the Red Sea, Black Sea and Panama Canal.](https://unctad.org/publication/navigating-troubled-waters-impact-global-trade-disruption-shipping-routes-red-sea-black) 5. [ 5 ][European Central Bank. (2024). Scenario analysis of a potential escalation of disruptions in the Red Sea area.](https://www.ecb.europa.eu/press/projections/html/ecb.projections202403_ecbstaff~f2f2d34d5a.en.html) 6. [ 6 ][World Bank. (2025). The deepening: MENA FCV Economic Series brief.](https://openknowledge.worldbank.org/bitstreams/de1a433c-e9ab-4eb3-ae7c-430c3a463d91/download) 7. [ 7 ][Board of Governors of the Federal Reserve System. (2023). Financial Stability Report: funding risks.](https://www.federalreserve.gov/publications/2023-may-financial-stability-report-funding-risks.htm) 8. [ 8 ][Office of Inspector General. (2023). Material loss review of Silicon Valley Bank.](https://oig.federalreserve.gov/reports/board-material-loss-review-silicon-valley-bank-sep2023.htm) 9. [ 9 ][World Health Organization. (2020). Novel coronavirus: China, disease outbreak news.](https://www.who.int/emergencies/disease-outbreak-news/item/2020-DON233) 10. [ 10 ][World Health Organization. (2020). Director-General's opening remarks at the media briefing on COVID-19, 11 March 2020.](https://www.who.int/news-room/speeches/item/who-director-general-s-opening-remarks-at-the-media-briefing-on-covid-19---11-march-2020) Continue exploring ## Complementary ontologies [237 categories GICS Industry Classification Join regional event migration to the industries operating or earning revenue in each location. Explore ontology →](https://nosible.com/ontologies/gics) [43 categories EM-DAT Disaster Classification Combine normalized location with disaster type to isolate regional physical-risk regimes. Explore ontology →](https://nosible.com/ontologies/em-dat) > Explore 231 geographic categories and example usages tracing Red Sea shipping risk, banking stress and pandemic reporting across regions and countries. **URL:** https://nosible.com/ontologies/geography --- --- title: "GICS Industry Classification" description: "Explore 237 GICS categories and example usages tracing AI concentration, semiconductor shortages and commercial-property stress across exposed industries." url: "https://nosible.com/ontologies/gics" --- [Home](https://nosible.com/) [Ontologies](https://nosible.com/ontologies) GICS Ontology field guide By NOSIBLE Research Updated 2026-07-18 # GICS Industry Classification GICS shows which parts of the economy an event concerns, preserving industry structure that a broad technology label removes. Industry exposure is not one variable. Semiconductor capacity, cloud-platform investment and data-centre buildout may all involve artificial intelligence, but they transmit through different companies, margins and risks. GICS separates them into sectors, groups, industries and sub-industries. World assigns that hierarchy to reporting so researchers can build cohorts at the required economic level. [[ 1 ]](https://www.msci.com/documents/1296102/11185224/GICS%20Methodology%202023.pdf) GICS supports peer selection, exposure measurement and spillover tests. It does not replace issuer fundamentals or a point-in-time security master. This guide provides the structure, World statistics, an AI-trade study and a falsifiable test of whether industry attention adds information beyond event volume. [[ 2 ]](https://doi.org/10.1046/j.1475-679X.2003.00122.x) Categories 237 categories Structure 4 levels World events labelled 23.0% Stable codes Since v1 Release data CC0 On this page [01 Foundations](https://nosible.com/ontologies/gics#foundations) [02 Categories](https://nosible.com/ontologies/gics#vocabulary) [03 Example usages](https://nosible.com/ontologies/gics#trends) [04 Downloads and references](https://nosible.com/ontologies/gics#downloads) Foundations ## GICS classification separates companies through four sector and industry levels globally GICS assigns companies according to their principal business activity. The hierarchy runs from sector to industry group, industry and sub-industry. Revenue is the primary classification input, with earnings and market perception used where business lines are mixed. The nested structure lets a researcher move from a broad economic exposure to a tighter operating peer group without changing classification systems. [[ 1 ]](https://www.msci.com/documents/1296102/11185224/GICS%20Methodology%202023.pdf) World classifies the economic activity expressed in an event record. It does not assert that every named company belongs to the selected sub-industry, and it does not resolve conglomerate revenue across segments. Event-level GICS labels should be joined to a point-in-time issuer classification when the test concerns security membership. Results should also survive alternative peer definitions because industry schemes differ in how well they explain returns, valuation multiples and growth. Event and issuer classifications answer different questions. [[ 2 ]](https://doi.org/10.1046/j.1475-679X.2003.00122.x) Categories ## GICS organises 237 categories into sectors, groups, industries and sub-industries GICS has four nested levels: 12 sector roots, followed by industry groups, industries and 164 leaf sub-industries. Each World assignment records one path at the available depth. The explorer exposes definitions, stable codes, full paths and World V1.2 counts. Search and bounded expansion keep the same interface usable across the complete hierarchy. [[ 1 ]](https://www.msci.com/documents/1296102/11185224/GICS%20Methodology%202023.pdf) Search the ontology 12 / 237 shown World V1.2 coverage Coverage **23.0%** Labelled events **3,516,971** World events **15,311,040** Ontology index 3,516,971 of 15,311,040 World events carry GICS Industry Classification labels. Select a category to inspect the evidence Hierarchy sector / industry_group / industry / sub_industry › Energy Oil, gas, drilling and field-service activity · 2 children › Materials · 5 children › Industrials · 3 children › Consumer Discretionary · 4 children › Consumer Staples · 3 children › Health Care · 2 children › Financials · 3 children › Information Technology · 3 children › Communication Services · 2 children › Utilities · 5 children › Real Estate · 2 children Other Energy ### Energy Copy link ↗ #### Definition Energy production, drilling and field services. Research use Oil, gas, drilling and field-service activity Code Energy Events with label 137,455 Share of labelled 3.9% Share of World 0.9% **Events with label** is the selected count. **Label share** divides it by 3,516,971 assigned events; **World share** divides it by all 15,311,040 events. #### Representative World V1.2 events One strong classified example per available year, with up to ten years shown. 10 examples 1. 2015-05-19Duke Energy Closes Aging Lake Julian Coal Plant in AshevilleCoverage 68 2. 2016-07-05Noble Energy Sells 3% Tamar Field Stake for $369 MillionCoverage 18 3. 2017-03-27Hurricane Energy Finds Largest Undeveloped UK Oil Field Off ShetlandCoverage 29 4. 2019-06-05Hurricane Energy Achieves First Oil at UK's Lancaster Fractured FieldCoverage 23 5. 2020-06-05Aker Energy Revives Ghana Pecan Offshore Oil Field Development PlansCoverage 21 6. 2021-11-23US Releases 50 Million Barrels Oil to Lower Energy CostsCoverage 1,461 7. 2022-06-06Hezbollah Threatens Israel Over Sea Boundary Energy DrillingCoverage 250 8. 2024-06-11Woodside Energy Achieves First Oil at Senegal's Sangomar Offshore FieldCoverage 183 9. 2025-05-14Valeura Energy Approves Wassana Field Redevelopment in ThailandCoverage 64 10. 2026-02-11Birchcliff Energy Reports Record 2025 Production and Strong Financial ResultsCoverage 132 Browse categories to compare definitions, World V1.2 statistics and representative classified events. Example usage 1 of 3 ### Example Usage: AI coverage rises sevenfold and concentrates in fewer industries over time World contains 13,730 entity-resolved English events that discuss artificial intelligence, carry a GICS sub-industry and name one of 15 compute or platform companies. Normalized coverage rises more than sevenfold after November 2022. The top-three sub-industry share increases from 48.7% to 59.9%, while effective sub-industries fall from 9.9 to 7.0. [[ 3 ]](https://openai.com/index/chatgpt/) Semiconductors rise from 19.8% to 28.7% of the cohort. A compute-infrastructure basket reaches 42.5%, up from 33.3%. Higher attention and narrower industry breadth mean an AI basket can conceal shared economic exposure. ChatGPT's launch sets the boundary; it does not establish causality or portfolio crowding. [[ 4 ]](https://www.sec.gov/Archives/edgar/data/1045810/000104581024000029/nvda-20240128.htm) [[ 5 ]](https://www.bis.gov/press-release/commerce-strengthens-restrictions-advanced-computing-semiconductors-semiconductor-manufacturing-equipment) Entity-resolved AI reporting · monthly 2020-01 to 2026-06 **AI-company coverage rises sevenfold and concentrates in three technology industries** stacked rate per million English events · three-month half-life Semiconductors Interactive Media & Services Application Software Other industries Before December 2022 **509 per million** December 2022 onward **3948 per million** The intensity line divides the remembered cohort by remembered English World events. Phase comparisons pool all events before December 2022 and all events afterwards. Rising normalized AI attention alongside fewer effective sub-industries indicates concentration in reporting exposure, not portfolio crowding or causality. Example usage 2 of 3 ### Example Usage: Automakers and chipmakers capture 84% of classified semiconductor-shortage coverage across exposed industries The 2021 semiconductor shortage is not only a chipmaker story. Automobile Manufacturers and Semiconductors each contribute 279 assignments, or 41.9% of the 666-event cohort. Together they account for 83.8% of classified coverage. Monthly normalized attention peaks at 2,198 events per million in October as production constraints persist. [[ 6 ]](https://www.commerce.gov/news/press-releases/2022/01/commerce-semiconductor-data-confirms-urgent-need-congress-pass-us) GICS separates suppliers from customers inside one supply-chain shock. Industry composition shows whether reporting concentrates on scarce components, curtailed vehicle production or downstream exposure. The result defines distinct equity cohorts for estimating inventory, margin and delivery effects instead of treating every semiconductor-shortage article as the same sector signal across portfolios in research. [[ 7 ]](https://www.whitehouse.gov/wp-content/uploads/2021/06/100-day-supply-chain-review-report.pdf) Automobile Manufacturers 41.9% Semiconductors 41.9% Other sub-industries 16.2% GICS sub-industry composition 666 assigned events **41.9 %** **41.9 %** **16.2 %** **Automobile Manufacturers** 279 events **Semiconductors** 279 events **Other sub-industries** 108 events The fixed English cohort requires semiconductor and supply-constraint terms. All 666 GICS-assigned events occur in 2021. Monthly counts are divided by English World volume for every month in the fixed year, enabling direct comparison throughout 2021. Example usage 3 of 3 ### Example Usage: Commercial-property attention reaches regional banks after office landlords peak during the credit cycle Office Real Estate Investment Trust attention reaches 238 events per million on 23 January 2024. New York Community Bancorp reports a quarterly loss on 31 January. Mortgage Finance then reaches 158 per million on 22 February, while Regional Banks peaks later at 205 per million on 28 March across the same credit cycle. [[ 8 ]](https://ir.flagstar.com/news-and-events/news-releases/press-release-details/2024/NEW-YORK-COMMUNITY-BANCORP-INC.-REPORTS-RECORD-RESULTS-FOR-2023/default.aspx) The sequence exposes transmission across balance sheets. Office landlords carry vacancy and valuation risk. Mortgage lenders and regional banks carry collateral, funding and credit risk. GICS supplies separate equity cohorts and event clocks for testing whether lender attention precedes revisions, deposits, spreads or relative returns. [[ 9 ]](https://www.federalreserve.gov/publications/2023-june-dodd-frank-act-stress-test-results.htm) Office REIT peak +238 per million Regional-bank peak +205 per million Mortgage-finance peak +158 per million Fifty-six-day normalized GICS sub-industry attention · per million Office REITs Regional Banks Mortgage Finance The fixed English cohort covers explicit commercial-property, office-vacancy and commercial-mortgage reporting from October 2022 through June 2024. Each line pools GICS sub-industry counts across the current and prior 55 days, then divides by all English World events in the same window. Data and sources ## Download GICS and References ### Downloads Download categories and counts as CSV, or the complete machine-readable release as JSON. [**CSV** Definitions and counts Download ↓](https://nosible.com/data/world-v1.2/ontologies/gics/v1/nodes.csv) [**JSON** Full machine-readable release Download ↓](https://nosible.com/data/world-v1.2/ontologies/gics/v1/statistics.json) Release details and citation files [Manifest](https://nosible.com/data/world-v1.2/ontologies/gics/v1/manifest.json) [Release README](https://nosible.com/data/world-v1.2/ontologies/gics/v1/README.md) [CC0 license and scope](https://nosible.com/data/world-v1.2/ontologies/gics/v1/LICENSE.md) [Citation file](https://nosible.com/data/world-v1.2/ontologies/gics/v1/CITATION.cff) [Changelog](https://nosible.com/data/world-v1.2/ontologies/gics/v1/CHANGELOG.md) Need the complete World event schema? [Open the World data dictionary.](https://nosible.com/data-dictionaries#world) ### References 1. [ 1 ][MSCI and S&P Dow Jones Indices. (2023). Global Industry Classification Standard (GICS) methodology.](https://www.msci.com/documents/1296102/11185224/GICS%20Methodology%202023.pdf) 2. [ 2 ][Bhojraj, S., Lee, C. M. C., & Oler, D. K. (2003). What's my line? A comparison of industry classification schemes for capital market research. Journal of Accounting Research, 41(5), 745–774.](https://doi.org/10.1046/j.1475-679X.2003.00122.x) 3. [ 3 ][OpenAI. (2022, November 30). Introducing ChatGPT.](https://openai.com/index/chatgpt/) 4. [ 4 ][NVIDIA Corporation. (2024). Form 10-K for the fiscal year ended January 28, 2024.](https://www.sec.gov/Archives/edgar/data/1045810/000104581024000029/nvda-20240128.htm) 5. [ 5 ][U.S. Bureau of Industry and Security. (2023, October 17). Commerce strengthens restrictions on advanced computing semiconductors and semiconductor manufacturing equipment.](https://www.bis.gov/press-release/commerce-strengthens-restrictions-advanced-computing-semiconductors-semiconductor-manufacturing-equipment) 6. [ 6 ][U.S. Department of Commerce. (2022). Semiconductor data confirms urgent need to address the chip shortage.](https://www.commerce.gov/news/press-releases/2022/01/commerce-semiconductor-data-confirms-urgent-need-congress-pass-us) 7. [ 7 ][The White House. (2021). Building resilient supply chains, revitalizing American manufacturing, and fostering broad-based growth: 100-day reviews.](https://www.whitehouse.gov/wp-content/uploads/2021/06/100-day-supply-chain-review-report.pdf) 8. [ 8 ][New York Community Bancorp. (2024). Full-year and fourth-quarter 2023 results.](https://ir.flagstar.com/news-and-events/news-releases/press-release-details/2024/NEW-YORK-COMMUNITY-BANCORP-INC.-REPORTS-RECORD-RESULTS-FOR-2023/default.aspx) 9. [ 9 ][Board of Governors of the Federal Reserve System. (2023). Dodd-Frank Act stress-test results.](https://www.federalreserve.gov/publications/2023-june-dodd-frank-act-stress-test-results.htm) Continue exploring ## Complementary ontologies [39 categories Asset Class Ontology Follow industry-level attention into the asset classes through which markets express the risk. Explore ontology →](https://nosible.com/ontologies/asset-classes) [416 categories NOSIBLE Event Ontology Separate the corporate actions driving an industry's changing event exposure. Explore ontology →](https://nosible.com/ontologies/nosible-events) > Explore 237 GICS categories and example usages tracing AI concentration, semiconductor shortages and commercial-property stress across exposed industries. **URL:** https://nosible.com/ontologies/gics --- --- title: "IAB Content Ontology" description: "Explore 709 IAB content categories and evidence showing how banking stress, labor disruption and hurricane risk rotate attention across investable subjects." url: "https://nosible.com/ontologies/iab-content-taxonomy" --- [Home](https://nosible.com/) [Ontologies](https://nosible.com/ontologies) IAB Content Ontology field guide By NOSIBLE Research Updated 2026-07-18 # IAB Content Ontology IAB Content separates a reported event's subject from its entities, industry and market sentiment through time for event research. Events about the same institution can concern different risks. A bank earnings release, depositor run, regulatory intervention and acquisition all name a bank but describe different subjects. IAB Content preserves those differences in a four-level hierarchy. World assigns one primary subject path to each canonical event. [[ 1 ]](https://iabtechlab.com/standards/content-taxonomy/) The hierarchy supports topic controls, thematic cohorts and tests of how attention changes during a market episode. It does not identify an issuer's industry or economic exposure. This guide provides every category, World base rates, a regional-bank failure study and downloadable evidence. Categories 709 categories Structure 4 levels World events labelled 100.0% Stable codes Since v1 Release data CC0 On this page [01 Foundations](https://nosible.com/ontologies/iab-content-taxonomy#foundations) [02 Categories](https://nosible.com/ontologies/iab-content-taxonomy#vocabulary) [03 Example usages](https://nosible.com/ontologies/iab-content-taxonomy#trends) [04 Downloads and references](https://nosible.com/ontologies/iab-content-taxonomy#downloads) Foundations ## IAB Content classifies subject matter, not companies, industries or economic exposure IAB Content provides a common hierarchy for subject matter across news, video, audio, games and other media. Broad roots support stable controls. Deeper categories separate narrower questions such as business banking, financial crises, consumer banking and financial reform. The standard was designed for contextual signalling and brand safety; research uses require independent validation against the target event cohort. [[ 1 ]](https://iabtechlab.com/standards/content-taxonomy/) World assigns the dominant subject expressed in each event record. It does not infer why the event matters to a security or whether investors acted on it. One primary path also compresses mixed subjects. Test neighbouring categories, source concentration and alternative text representations before treating a topic shift as an economic regime change. Categories ## IAB Content uses four levels to map 709 distinct subjects The hierarchy contains 39 top-level subjects and 619 leaves across four tiers. Each World event carries one path at the available depth. The explorer exposes definitions, stable codes, full paths and World V1.2 counts. Search and bounded expansion keep all 709 categories accessible on this page without fragmenting the reference across separate pages. [[ 1 ]](https://iabtechlab.com/standards/content-taxonomy/) Search the ontology 39 / 709 shown World V1.2 coverage Coverage **100.0%** Labelled events **15,311,040** World events **15,311,040** Ontology index 15,311,040 of 15,311,040 World events carry IAB Content Ontology labels. Select a category to inspect the evidence Hierarchy tier_1 / tier_2 / tier_3 / tier_4 › Attractions Visitor destinations and leisure venues · 11 children › Automotive · 16 children › Books and Literature · 4 children › Business and Finance Corporate activity, industries, capital and market structure · 3 children › Careers · 6 children Communication Crime Disasters › Education · 12 children › Entertainment · 29 children › Events · 3 children › Family and Relationships · 7 children › Fine Art · 8 children › Food & Drink · 12 children › Genres · 29 children › Healthy Living · 8 children › Hobbies & Interests · 16 children Attractions ### Attractions Copy link ↗ #### Definition Entertainment and leisure destinations that attract visitors, including landmarks, amusement venues and recreational sites. Research use Visitor destinations and leisure venues Code Attractions Events with label 200,167 Share of labelled 1.3% Share of World 1.3% **Events with label** is the selected count. **Label share** divides it by 15,311,040 assigned events; **World share** divides it by all 15,311,040 events. #### Representative World V1.2 events One strong classified example per available year, with up to ten years shown. 10 examples 1. 2015-05-07Nintendo Partners With Universal For New Theme Park AttractionsCoverage 213 2. 2016-04-20Disney Announces Toy Story Land Attractions Including Slinky Dog DashCoverage 13 3. 2017-06-13UK Terror Attacks Cause Visitor Numbers to Drop at London AttractionsCoverage 18 4. 2019-08-26Disney Announces New Mary Poppins and Wakanda Attractions at EpcotCoverage 113 5. 2020-03-18Coronavirus Closures Hit Global Attractions and Las Vegas ShowsCoverage 109 6. 2021-11-20Disney Unveils New Park Attractions and Lore at Destination D23 2021Coverage 25 7. 2022-11-02Universal Studios Florida Closes Five Kids Attractions for New Family EntertainmentCoverage 165 8. 2024-08-02Arizona showcases quirky attractions from The Thing to a new poop museumCoverage 76 9. 2025-08-20Museum of Ice Cream Singapore Unveils New Attractions and FlavorsCoverage 107 10. 2026-05-21UK Cuts VAT on Summer Attractions to Boost Cost of LivingCoverage 362 Browse categories to compare definitions, World V1.2 statistics and representative classified events. Example usage 1 of 3 ### Example Usage: Regional-bank coverage rotates from routine banking into crisis within just days World contains 512 English events naming Silicon Valley Bank, Signature Bank or First Republic Bank from January through June 2023. The banks failed on 10 March, 12 March and 1 May. Daily coverage is divided by all English World events and given a seven-day half-life, separating the shock from overall corpus growth. [[ 2 ]](https://www.fdic.gov/news/press-releases/2023/pr23019.html) [[ 3 ]](https://www.dfs.ny.gov/consumers/alerts/SignatureBank) [[ 4 ]](https://www.fdic.gov/news/press-releases/2023/pr23034.html) The Business and Finance root stays near 90%, but its leaves rotate. Business Banking and Finance falls from 63.2% before failure to 11.2% during intervention. Financial Crisis rises from zero to 41.8%. That transition separates routine bank coverage from systemic-stress reporting and creates a cleaner clock for peer-contagion, deposit and policy tests. [[ 5 ]](https://www.federalreserve.gov/publications/review-of-the-federal-reserves-supervision-and-regulation-of-silicon-valley-bank.htm) Entity-resolved regional-bank cohort 512 assigned events · 2023-01-01 to 2023-06-30 **Routine banking coverage becomes crisis reporting within days** 7-day decayed share of assigned events Routine banking Financial crisis Policy and depositors The fixed cohort names Silicon Valley Bank, Signature Bank or First Republic Bank. Lines apply a seven-day half-life to assigned category counts. “Policy and depositors” combines Financial Reform and Consumer Banking. The chart describes classified reporting, not deposits, contagion or returns. Example usage 2 of 3 ### Example Usage: Hollywood strikes move coverage from scripts into production, studios and distribution Hollywood strike coverage changes category when actors join writers on 14 July 2023. Screenwriting falls from 39.2% to 11.4% of assigned IAB events. Movies rises from 8.2% to 21.4%, while Entertainment Industry rises from 5.7% to 15.5%. The reporting object moves from scripts into production and distribution disruption. [[ 9 ]](https://www.sagaftra.org/sites/default/files/sa_documents/Strike%20Notice%20to%20Members.pdf) The category shift defines two different exposure windows. The writers-only phase concentrates on content creation and contract terms. The joint strike extends the shock to performers, release schedules, studios and streaming platforms. An entertainment keyword cannot separate those mechanisms; IAB categories identify when the affected system expands beyond the original labor dispute and across the wider commercial value chain. [[ 10 ]](https://www.wga.org/uploadedfiles/the-guild/annual-report/annualreport23.pdf) Screenwriting -27.8pp Movies +13.2pp Entertainment Industry +9.8pp Fourteen-day pooled share of IAB-assigned Hollywood strike events · % Screenwriting Movies Entertainment Industry The fixed English strike cohort contains 1,264 IAB-assigned events from May through November 2023. Each line pools category counts across the current and prior thirteen days, then divides by all assigned cohort events. Markers identify the writers' strike, actors' strike and writers' agreement. Example usage 3 of 3 ### Example Usage: Home Insurance enters Hurricane Ian coverage after Florida landfall exposes household financial risk Hurricane Ian converts a weather cohort into an insurance cohort after Florida landfall on 28 September 2022. Normalized IAB-assigned coverage rises from 2,371 to 9,257 events per million across matched seven-day windows, a 3.9-fold increase after controlling for corpus growth. Disaster assignments increase from 26.3% to 34.8%. [[ 7 ]](https://www.nhc.noaa.gov/data/tcr/AL092022_Ian.pdf) Home Insurance rises from 2.6% before landfall to 5.5% afterwards. The shift identifies when physical hazard becomes household balance-sheet exposure. Weather keywords capture the storm; IAB categories separate property risk, insurance demand and wider consumer consequences after damage assessments and claims arrive. [[ 8 ]](https://www.iii.org/press-release/triple-i-ian-brings-financial-first-response-to-fore-begin-claims-filing-093022) Normalized coverage +6,886 per million Disasters +8.5pp Home Insurance +2.9pp Weather -2.7pp Seven-day pooled share of IAB-assigned events · % Home Insurance Disasters Weather The explicit English Hurricane Ian cohort compares the seven days before 28 September 2022 with the period from landfall. Each line pools category counts across the current and prior six days, then divides by all IAB-assigned cohort events in that window. Data and sources ## Download IAB Content and References ### Downloads Download categories and counts as CSV, or the complete machine-readable release as JSON. [**CSV** Definitions and counts Download ↓](https://nosible.com/data/world-v1.2/ontologies/iab-content-taxonomy/v1/nodes.csv) [**JSON** Full machine-readable release Download ↓](https://nosible.com/data/world-v1.2/ontologies/iab-content-taxonomy/v1/statistics.json) Release details and citation files [Manifest](https://nosible.com/data/world-v1.2/ontologies/iab-content-taxonomy/v1/manifest.json) [Release README](https://nosible.com/data/world-v1.2/ontologies/iab-content-taxonomy/v1/README.md) [CC0 license and scope](https://nosible.com/data/world-v1.2/ontologies/iab-content-taxonomy/v1/LICENSE.md) [Citation file](https://nosible.com/data/world-v1.2/ontologies/iab-content-taxonomy/v1/CITATION.cff) [Changelog](https://nosible.com/data/world-v1.2/ontologies/iab-content-taxonomy/v1/CHANGELOG.md) Need the complete World event schema? [Open the World data dictionary.](https://nosible.com/data-dictionaries#world) ### References 1. [ 1 ][IAB Tech Lab. (2024). Content Ontology 3.0.](https://iabtechlab.com/standards/content-taxonomy/) 2. [ 2 ][Federal Deposit Insurance Corporation. (2023, March 13). FDIC acts to protect all depositors of the former Silicon Valley Bank.](https://www.fdic.gov/news/press-releases/2023/pr23019.html) 3. [ 3 ][New York State Department of Financial Services. (2023, March 12). Notice regarding Signature Bank.](https://www.dfs.ny.gov/consumers/alerts/SignatureBank) 4. [ 4 ][Federal Deposit Insurance Corporation. (2023, May 1). JPMorgan Chase assumes all deposits of First Republic Bank.](https://www.fdic.gov/news/press-releases/2023/pr23034.html) 5. [ 5 ][Board of Governors of the Federal Reserve System. (2023). Review of the Federal Reserve's supervision and regulation of Silicon Valley Bank.](https://www.federalreserve.gov/publications/review-of-the-federal-reserves-supervision-and-regulation-of-silicon-valley-bank.htm) 6. [ 6 ][Jiang, E. X., Matvos, G., Piskorski, T., & Seru, A. (2024). Monetary tightening and U.S. bank fragility in 2023. Journal of Financial Economics, 159, 103899.](https://doi.org/10.3386/w31048) 7. [ 9 ][SAG-AFTRA. (2023). TV/Theatrical/Streaming strike notice and order, effective 14 July 2023.](https://www.sagaftra.org/sites/default/files/sa_documents/Strike%20Notice%20to%20Members.pdf) 8. [ 10 ][Writers Guild of America West. (2023). Annual report documenting the work stoppage that began 2 May 2023.](https://www.wga.org/uploadedfiles/the-guild/annual-report/annualreport23.pdf) 9. [ 7 ][National Hurricane Center. (2023). Tropical Cyclone Report: Hurricane Ian, 23-30 September 2022.](https://www.nhc.noaa.gov/data/tcr/AL092022_Ian.pdf) 10. [ 8 ][Insurance Information Institute. (2022). Ian brings financial first response to fore: begin claims-filing.](https://www.iii.org/press-release/triple-i-ian-brings-financial-first-response-to-fore-begin-claims-filing-093022) Continue exploring ## Complementary ontologies [718 categories IPTC Media Topics Compare web-content categories with a news-native subject hierarchy to test classification stability. Explore ontology →](https://nosible.com/ontologies/iptc-media-topics) [237 categories GICS Industry Classification Translate subject attention into the industries that may carry the economic exposure. Explore ontology →](https://nosible.com/ontologies/gics) > Explore 709 IAB content categories and evidence showing how banking stress, labor disruption and hurricane risk rotate attention across investable subjects. **URL:** https://nosible.com/ontologies/iab-content-taxonomy --- --- title: "ICD-11 Chapters and Blocks" description: "Explore 300 ICD-11 chapter and block categories with evidence on H5N1, GLP-1 and mental-health reporting across distinct clinical and market contexts." url: "https://nosible.com/ontologies/icd-11" --- [Home](https://nosible.com/) [Ontologies](https://nosible.com/ontologies) ICD-11 Ontology field guide By NOSIBLE Research Updated 2026-07-18 # ICD-11 Chapters and Blocks ICD-11 shows when a named health risk requires more detail than a broad disease chapter provides for exposure analysis through time. Health reporting can remain in a broad chapter or gain qualifiers. ICD-11 preserves that distinction. World joins assignments to entities, events and source-time metadata without treating news as a clinical record. [[ 1 ]](https://www.who.int/news-room/fact-sheets/detail/icd-11) This guide tracks reporting from broad disease chapters into Extension Codes. The H5N1 study starts with the March 2024 US dairy detection. It measures context, not infections, transmission or losses. [[ 2 ]](https://www.aphis.usda.gov/news/agency-announcements/usda-hhs-announce-new-actions-reduce-impact-spread-h5n1) [[ 4 ]](https://doi.org/10.1038/s41586-024-07849-4) Categories 300 categories Structure 2 levels World events labelled 7.8% Stable codes Since v1 Release data CC0 On this page [01 Foundations](https://nosible.com/ontologies/icd-11#foundations) [02 Categories](https://nosible.com/ontologies/icd-11#vocabulary) [03 Example usages](https://nosible.com/ontologies/icd-11#trends) [04 Downloads and references](https://nosible.com/ontologies/icd-11#downloads) Foundations ## ICD-11 classifies reported health conditions without measuring incidence, transmission or severity The World Health Organization created ICD-11 for consistent recording, reporting and comparison of health conditions. The full standard contains diagnostic categories, extension codes and combinable concepts. World uses 28 chapters and 272 broad blocks. That level supports event research without implying clinical precision. [[ 1 ]](https://www.who.int/news-room/fact-sheets/detail/icd-11) World classifies the health subject expressed in an event. It does not diagnose a person, infer incidence or establish treatment efficacy. One event can discuss a disease, regulator, company and affected population while receiving one chapter and block path. Counts must be checked for repeated coverage, source concentration and assignment error. [[ 1 ]](https://www.who.int/news-room/fact-sheets/detail/icd-11) Categories ## ICD-11 organises 300 health labels across broad chapters and specific diagnostic blocks The ontology has two levels. Chapters define broad health systems. Blocks provide the finest World V1.2 classification. The explorer shows every definition, code, path and corpus count. Detailed disease questions still require an explicit text or entity cohort because a broad block does not identify every named pathogen. [[ 1 ]](https://www.who.int/news-room/fact-sheets/detail/icd-11) Search the ontology 28 / 300 shown World V1.2 coverage Coverage **7.8%** Labelled events **1,194,981** World events **15,311,040** Ontology index 1,194,981 of 15,311,040 World events carry ICD-11 Chapters and Blocks labels. Select a category to inspect the evidence Hierarchy chapter / block › Certain infectious or parasitic diseases Outbreaks, vaccines and health responses · 21 children › Neoplasms Cancer incidence, treatment, trial and regulatory cohorts · 7 children › Diseases of the blood or blood-forming organs · 3 children › Diseases of the immune system · 6 children › Endocrine, nutritional or metabolic diseases Diabetes, obesity, metabolic treatment and nutrition cohorts · 4 children › Mental, behavioural or neurodevelopmental disorders Mental-health demand, treatment and policy cohorts · 20 children › Sleep-wake disorders · 6 children › Diseases of the nervous system Neurology, degeneration and central-nervous-system treatment cohorts · 18 children › Diseases of the visual system · 11 children › Diseases of the ear or mastoid process · 6 children › Diseases of the circulatory system Cardiovascular risk, outcomes and treatment cohorts · 15 children › Diseases of the respiratory system · 7 children › Diseases of the digestive system · 17 children › Diseases of the skin · 14 children › Diseases of the musculoskeletal system or connective tissue · 4 children › Diseases of the genitourinary system · 6 children › Conditions related to sexual health · 3 children Certain infectious or parasitic diseases ### Certain infectious or parasitic diseases Copy link ↗ #### Definition Diseases caused by infectious or parasitic agents, including related outbreaks and public-health responses. Research use Outbreaks, vaccines and health responses Code Certain infectious or parasitic diseases Events with label 100,636 Share of labelled 8.4% Share of World 0.7% **Events with label** is the selected count. **Label share** divides it by 1,194,981 assigned events; **World share** divides it by all 15,311,040 events. #### Representative World V1.2 events One strong classified example per available year, with up to ten years shown. 10 examples 1. 2015-02-17Climate Change Drives Rapid Emergence of Infectious Diseases GloballyCoverage 32 2. 2016-04-11Human Diseases from Africa May Have Caused Neanderthal ExtinctionCoverage 23 3. 2017-04-19World Leaders Pledge 812 Million to End Neglected Tropical DiseasesCoverage 42 4. 2019-12-27US Measles Outbreaks Surge in 2019 Amid Global Infectious Disease CrisisCoverage 16 5. 2020-05-14Nigerian Governors Urge National Assembly to Halt Infectious Diseases BillCoverage 31 6. 2021-06-10NSW and Queensland on High Alert After Victorian Couple Travels InfectiousCoverage 40 7. 2022-08-08Study: Climate Change Worsens 58% of Known Human Infectious DiseasesCoverage 155 8. 2024-10-30WHO: Tuberculosis Surpasses COVID-19 as Top Infectious Disease KillerCoverage 190 9. 2025-11-02ICMR Study Finds One in Nine Indians Tested Positive for Infectious DiseasesCoverage 27 10. 2026-04-02CDC Halts Dozens of Infectious Disease Tests Amid Staffing CutsCoverage 141 Browse categories to compare definitions, World V1.2 statistics and representative classified events. Example usage 1 of 3 ### Example Usage: H5N1 coverage gains investable context after the virus reaches United States dairy herds US authorities confirmed H5N1 in dairy cattle on 25 March 2024 and a Texas dairy worker on 1 April. Exact H5N1 attention rises 7.53 times, from 4.0 to 29.8 events per 10,000 relevant health events. Infectious Diseases falls from 69.2% to 48.9% of the assigned cohort. Extension Codes rises from 30.8% to 51.1%. [[ 2 ]](https://www.aphis.usda.gov/news/agency-announcements/usda-hhs-announce-new-actions-reduce-impact-spread-h5n1) [[ 3 ]](https://www.cdc.gov/mmwr/volumes/73/wr/mm7321e1.htm) [[ 4 ]](https://doi.org/10.1038/s41586-024-07849-4) Infectious Diseases supplies the broad chapter. Extension Codes add qualifiers for exposure triage. Their share rises from 30.8% to 51.1% after dairy detection, identifying a specific reporting cohort for workers, producers and supply chains. H5N1 language defines the cohort; it does not measure infections, severity, transmission or economic loss. [[ 5 ]](https://www.fsa.usda.gov/news-events/news/07-01-2024/usda-begin-accepting-applications-expanded-emergency-livestock) [[ 6 ]](https://doi.org/10.1038/s43247-025-03153-9) [[ 7 ]](https://www.cdc.gov/media/releases/2025/m0106-h5-birdflu-death.html) Explicit H5N1 cohort · complete weeks 539 events · 2022-01-03 to 2026-06-28 **Normalised H5N1 attention rises after US dairy detection** 28-day rate per 10,000 health events Infectious Diseases is the broad chapter. Extension Codes carry supplemental context used to qualify a health condition. Their rise does not measure outbreak severity or supply-chain loss. It identifies a more specific reporting cohort that researchers can join to affected entities, locations and agricultural exposures. The attention chart uses a trailing 28-day rate and preserves the health-event denominator. Example usage 2 of 3 ### Example Usage: GLP-1 coverage expands from metabolic treatment into cardiovascular and adverse-event contexts GLP-1 reporting moves beyond diabetes and obesity as clinical evidence and approvals expand. Metabolic classifications remain central, while cardiovascular, digestive and adverse-event contexts produce distinct bursts of normalized attention. The chart aligns those category-specific rates with major trial and approval dates rather than pooling each Ozempic, Wegovy or Zepbound event inside one undifferentiated clinical series. [[ 8 ]](https://doi.org/10.1056/NEJMoa2307563) The distinction changes the exposure map. Metabolic events address obesity and diabetes demand. Cardiovascular outcomes expand the eligible population and reimbursement case. Digestive and injury classifications surface safety risk. Researchers can connect each clinical channel to manufacturers, payers, providers and consumer sectors while preserving a separate event clock for each investment thesis. [[ 9 ]](https://www.fda.gov/news-events/press-announcements/fda-approves-first-treatment-reduce-risk-serious-heart-problems-specifically-adults-obesity-or) Endocrine and metabolic +11.3pp Extension Codes +28.3pp Circulatory system +3.6pp Digestive system +4.5pp Twenty-one-day normalized category attention · per million Metabolic Cardiovascular Digestive Adverse events The fixed English cohort covers 11 July 2023 through 20 March 2024 and requires a named GLP-1 drug or GLP-1 term. Each line pools category counts across the current and prior twenty days, then divides by all English World events in that window. Markers identify public trial or regulatory milestones. Example usage 3 of 3 ### Example Usage: Mental-health reporting shifts from mortality toward clinical disorders during the first lockdowns During the first COVID-19 lockdowns, the clinical meaning of mental-health coverage changes within weeks. Mental, Behavioral or Neurodevelopmental Disorders rises from 34.9% of ICD-assigned events in the 28 days ending 10 March to 68.4% by 21 April. External Causes of Morbidity or Mortality falls from 49.5% to 16.5%. [[ 10 ]](https://www.who.int/publications/i/item/WHO-2019-nCoV-MentalHealth-2020.1) That rotation separates clinical demand from mortality reporting. The first channel maps to care utilization, telehealth, pharmaceuticals, insurers, employers and productivity. The second maps more directly to suicide and public-safety outcomes. ICD-11 supplies separate event clocks for testing demand, claims, workforce disruption and policy responses without treating mental-health reporting as one signal in model testing. [[ 11 ]](https://www.oecd.org/coronavirus/policy-responses/tackling-the-mental-health-impact-of-the-covid-19-crisis-an-integrated-whole-of-society-response-0ccafa0b/) Clinical mental health +33.5pp External causes -33.0pp Twenty-eight-day pooled share of ICD-assigned mental-health events · % Clinical mental health External causes The fixed English cohort requires mental-health, depression, anxiety-disorder or suicide language from January through September 2020. Each line pools ICD chapter counts across the current and prior twenty-seven days, then divides by all ICD-assigned cohort events. It measures classified reporting, not clinical prevalence, diagnosis rates or population-level changes in mental-health outcomes. Data and sources ## Download ICD-11 and References ### Downloads Download categories and counts as CSV, or the complete machine-readable release as JSON. [**CSV** Definitions and counts Download ↓](https://nosible.com/data/world-v1.2/ontologies/icd-11/v1/nodes.csv) [**JSON** Full machine-readable release Download ↓](https://nosible.com/data/world-v1.2/ontologies/icd-11/v1/statistics.json) Release details and citation files [Manifest](https://nosible.com/data/world-v1.2/ontologies/icd-11/v1/manifest.json) [Release README](https://nosible.com/data/world-v1.2/ontologies/icd-11/v1/README.md) [CC0 license and scope](https://nosible.com/data/world-v1.2/ontologies/icd-11/v1/LICENSE.md) [Citation file](https://nosible.com/data/world-v1.2/ontologies/icd-11/v1/CITATION.cff) [Changelog](https://nosible.com/data/world-v1.2/ontologies/icd-11/v1/CHANGELOG.md) Need the complete World event schema? [Open the World data dictionary.](https://nosible.com/data-dictionaries#world) ### References 1. [ 1 ][World Health Organization. (2022). ICD-11 fact sheet.](https://www.who.int/news-room/fact-sheets/detail/icd-11) 2. [ 2 ][US Department of Agriculture. (2024). USDA and HHS announce new actions to reduce the impact and spread of H5N1.](https://www.aphis.usda.gov/news/agency-announcements/usda-hhs-announce-new-actions-reduce-impact-spread-h5n1) 3. [ 3 ][Uyeki, T. M. et al. (2024). Highly pathogenic avian influenza A(H5N1) virus infection in a dairy farm worker. Morbidity and Mortality Weekly Report, 73, 501-505.](https://www.cdc.gov/mmwr/volumes/73/wr/mm7321e1.htm) 4. [ 4 ][Caserta, L. C. et al. (2024). Spillover of highly pathogenic avian influenza H5N1 virus to dairy cattle. Nature, 634, 669-676.](https://doi.org/10.1038/s41586-024-07849-4) 5. [ 5 ][US Department of Agriculture. (2024, July 1). USDA begins accepting applications for expanded Emergency Assistance for Livestock, Honeybees, and Farm-raised Fish Program assistance.](https://www.fsa.usda.gov/news-events/news/07-01-2024/usda-begin-accepting-applications-expanded-emergency-livestock) 6. [ 6 ][Morel, C. M. et al. (2026). The economic burden of highly pathogenic avian influenza H5N1. Communications Earth & Environment.](https://doi.org/10.1038/s43247-025-03153-9) 7. [ 7 ][US Centers for Disease Control and Prevention. (2025, January 6). CDC reports first H5 bird flu death in United States.](https://www.cdc.gov/media/releases/2025/m0106-h5-birdflu-death.html) 8. [ 8 ][Lincoff, A. M. et al. (2023). Semaglutide and cardiovascular outcomes in obesity without diabetes. New England Journal of Medicine, 389, 2221–2232.](https://doi.org/10.1056/NEJMoa2307563) 9. [ 9 ][U.S. Food and Drug Administration. (2024). FDA approves first treatment to reduce serious cardiovascular risks in adults with obesity or overweight.](https://www.fda.gov/news-events/press-announcements/fda-approves-first-treatment-reduce-risk-serious-heart-problems-specifically-adults-obesity-or) 10. [ 10 ][World Health Organization. (2020). Mental health and psychosocial considerations during the COVID-19 outbreak.](https://www.who.int/publications/i/item/WHO-2019-nCoV-MentalHealth-2020.1) 11. [ 11 ][Organisation for Economic Co-operation and Development. (2021). Tackling the mental health impact of the COVID-19 crisis.](https://www.oecd.org/coronavirus/policy-responses/tackling-the-mental-health-impact-of-the-covid-19-crisis-an-integrated-whole-of-society-response-0ccafa0b/) Continue exploring ## Complementary ontologies [237 categories GICS Industry Classification Map clinical event regimes to the healthcare and consumer industries carrying financial exposure. Explore ontology →](https://nosible.com/ontologies/gics) [186 categories UN Sustainable Development Goals Connect health classifications to the development targets they affect across countries and time. Explore ontology →](https://nosible.com/ontologies/sustainable-development-goals) > Explore 300 ICD-11 chapter and block categories with evidence on H5N1, GLP-1 and mental-health reporting across distinct clinical and market contexts. **URL:** https://nosible.com/ontologies/icd-11 --- --- title: "IPTC News Genre" description: "Explore 55 IPTC News Genre categories, World V1.2 evidence and example usages showing how forecasts, obituaries and supplied formats change event research." url: "https://nosible.com/ontologies/iptc-genre" --- [Home](https://nosible.com/) [Ontologies](https://nosible.com/ontologies) IPTC Genre Ontology field guide By NOSIBLE Research Updated 2026-07-18 # IPTC News Genre IPTC News Genre separates expectations from official statements and reported outcomes throughout each scheduled market decision cycle. Two articles can describe the same event but serve different purposes. A forecast states expectations. A curtain raiser prepares readers for a scheduled event. A press release carries an institution’s statement. Results report an outcome. Analysis interprets it. Genre preserves that role across subjects, sentiment and sources. [[ 1 ]](https://www.iptc.org/std/NewsCodes/treeview/genre/genre-en-GB.html) The distinction matters around scheduled market events. Expectations form before a decision, official information arrives at a known time, and reporting then resolves uncertainty. Genre separates those stages before tests of tone, novelty or market response. World assigns one of 55 peer labels to every canonical event, making information state observable through event time. That sequence is directly testable. [[ 2 ]](https://www.federalreserve.gov/monetarypolicy/fomccalendars.htm) [[ 3 ]](https://doi.org/10.1111/jofi.12196) [[ 4 ]](https://doi.org/10.1093/qje/qjy004) Categories 55 categories Structure Flat, no hierarchy World events labelled 100.0% Stable codes Since v1 Release data CC0 On this page [01 Foundations](https://nosible.com/ontologies/iptc-genre#foundations) [02 Categories](https://nosible.com/ontologies/iptc-genre#vocabulary) [03 Example usages](https://nosible.com/ontologies/iptc-genre#trends) [04 Downloads and references](https://nosible.com/ontologies/iptc-genre#downloads) Foundations ## News genre identifies editorial purpose without judging accuracy, quality or market impact IPTC News Genre classifies editorial form. Its labels distinguish current events, forecasts, interviews, press releases, transcripts, results, analysis and other publishing conventions. The label does not describe the event’s subject or economic importance. It tells a researcher how the information was presented and what role the item was designed to perform. [[ 1 ]](https://www.iptc.org/std/NewsCodes/treeview/genre/genre-en-GB.html) World assigns the dominant genre of the canonical event record. It does not verify factual accuracy, identify the original publisher’s intent or measure whether investors read the item. Genre can separate expectation-setting from outcome reporting, but a valid market test must still control for content, source, timing, novelty and policy surprise. [[ 5 ]](https://doi.org/10.1080/08997764.2018.1515767) [[ 6 ]](https://doi.org/10.17016/FEDS.2025.048) Categories ## IPTC News Genre uses 55 flat categories for editorial form The ontology is flat. Every event receives one peer-level label, with no parent or child classes. The explorer exposes the formal definition, stable code and World V1.2 frequency for all 55 labels. Search keeps the interface consistent with larger hierarchical ontologies and preserves explicit category boundaries. Search the ontology 55 / 55 shown World V1.2 coverage Coverage **100.0%** Labelled events **15,311,040** World events **15,311,040** Ontology index 15,311,040 of 15,311,040 World events carry IPTC News Genre labels. Select a category to inspect the evidence Categories iptc_genre Actuality Actuality captures point-in-time observations as events unfold, separating live reporting from later interpretation Advertiser Supplied Advice Advisory Analysis Interpretation of causes, implications and context Anniversary Archival material Background Behind the Story Biography Birth Announcement Current Events Immediate reporting of a developing event Curtain Raiser Preparation for a known future event Daybook Exclusive Fact Check Feature Actuality ### Actuality Copy link ↗ #### Definition Factual, time-stamped reporting that records an event's chronology, observations and immediate developments as they occur. Research use Actuality captures point-in-time observations as events unfold, separating live reporting from later interpretation Code Actuality Events with label 151,919 Share of labelled 1.0% Share of World 1.0% **Events with label** is the selected count. **Label share** divides it by 15,311,040 assigned events; **World share** divides it by all 15,311,040 events. #### Representative World V1.2 events One strong classified example per available year, with up to ten years shown. 10 examples 1. 2015-10-19US Astronaut Scott Kelly Breaks Record for Most Time in SpaceCoverage 16 2. 2016-02-29Google Self-Driving Car Hits Bus in California for First TimeCoverage 91 3. 2017-08-04Robert Pattinson Refuses Dog Sex Scene in Good Time FilmCoverage 68 4. 2019-02-25Rami Malek Wins Best Actor Oscar, Falls Off Stage During Acceptance SpeechCoverage 7,920 5. 2020-05-21NASA Scientists Claim Evidence of Parallel Universe With Backward Time FlowCoverage 37 6. 2021-10-22Alec Baldwin Kills Cinematographer Halyna Hutchins With Prop Gun on Rust SetCoverage 6,644 7. 2022-12-15Physicists Struggle to Define Time and RealityCoverage 111 8. 2024-07-10We Live in Time Trailer Reveals Romance Between Garfield and PughCoverage 264 9. 2025-06-26Faith Kipyegon Fails Sub-4 Mile Attempt in Paris Despite Fastest Women's TimeCoverage 460 10. 2026-04-07Artemis II Crew Breaks Apollo 13 Distance Record During Lunar FlybyCoverage 12,081 Browse categories to compare definitions, World V1.2 statistics and representative classified events. Example usage 1 of 3 ### Example Usage: FOMC genres separate market expectations from confirmed outcomes around scheduled policy decisions Across 40 Federal Open Market Committee decisions from 2021 to 2025, Forecast and Curtain Raiser form 56.5% of pre-announcement coverage. Their share falls after each decision as Results and Press Release expand, without becoming a results-led regime. Matched weekdays confirm a smaller information clock around scheduled policy decisions. [[ 2 ]](https://www.federalreserve.gov/monetarypolicy/fomccalendars.htm) That clock prevents look-ahead leakage and cohort contamination. Pre-decision articles encode expectations; post-decision articles contain outcomes and official communication. Pooling them can make a model appear predictive when it has learned information released after the tradeable timestamp. Genre labels let researchers isolate anticipation, announcement and digestion before testing rates, currencies, equities or volatility. [[ 2 ]](https://www.federalreserve.gov/monetarypolicy/fomccalendars.htm) **FOMC coverage moves from expectations toward outcomes after each decision** forward-looking share minus resolved share Meeting window Matched window The coverage chart divides Federal Reserve events by all World events on the same dates. The forward-looking share combines Forecast and Curtain Raiser. The resolved share combines Results Listings and Statistics with Press Release. Each point pools 40 scheduled decisions. Matched dates use the same weekdays two weeks earlier. The comparison is descriptive and does not measure policy surprise or market impact. Example usage 2 of 3 ### Example Usage: Obituary labels remove recycled history from event studies around major deaths Obituaries account for 59.9% of 554 Queen Elizabeth II events from her death through the funeral, compared with 3.7% of the same-date English baseline. The classification identifies the dominant block of retrospective reporting inside a global news surge rather than allowing it to masquerade as hundreds of independent information arrivals. [[ 7 ]](https://www.gov.uk/government/speeches/prime-ministers-statement-on-the-death-of-her-majesty-queen-elizabeth-ii) This matters whenever a prominent death affects currencies, sovereign institutions, succession risk or exposed companies. Obituaries frequently repeat known history. Removing or down-weighting them prevents duplicated biography from inflating attention, sentiment and novelty measures. Researchers retain live announcements and institutional consequences while applying a cleaner event clock to genuinely new information. [[ 1 ]](https://www.iptc.org/std/NewsCodes/treeview/genre/genre-en-GB.html) Obituary +56.2pp Composition of classified events share of each cohort **English baseline** 96.3 % **Queen cohort** 59.9 % 40.1 % Obituary Other genres The fixed cohort contains 554 events from 8 to 22 September 2022, equal to 19,129 per million English World events. The same-date baseline contains 28,961 events during the matched two-week period surrounding the funeral and succession. Example usage 3 of 3 ### Example Usage: iPhone 16 coverage moves from previews into supplied and results-led formats iPhone 16 launch coverage rises 69% after Apple's 9 September announcement, from 1,649 to 2,785 events per million. Curtain Raiser falls from 40.8% to 14.9% of assigned genres. Advertiser Supplied rises from 2.9% to 14.4%, while Results Listings and Statistics enters the reporting mix at 7.9% after launch. [[ 8 ]](https://www.apple.com/newsroom/2024/09/apple-introduces-iphone-16-and-iphone-16-plus/) The transition exposes both an event clock and a source-format risk. Preview material arrives before the launch; supplied and results-led formats expand afterwards. Researchers can separate anticipation from product disclosure, but Advertiser Supplied content requires source controls before interpreting tone, novelty or market impact. Genre classification makes that contamination visible rather than silently pooling it inside one cohort. [[ 1 ]](https://www.iptc.org/std/NewsCodes/treeview/genre/genre-en-GB.html) Curtain Raiser -25.9pp Advertiser Supplied +11.5pp Results +7.9pp Normalized coverage +1,136 per million Seven-day pooled share of assigned genres · % Advertiser Supplied Curtain Raiser Results The fixed English cohort covers 15 August through 3 October 2024 and requires iPhone 16 or Apple Intelligence. Each line pools genre counts across the current and prior six days, then divides by all genre-assigned cohort events in that window. The marker is Apple's 9 September announcement. Data and sources ## Download IPTC Genre and References ### Downloads Download categories and counts as CSV, or the complete machine-readable release as JSON. [**CSV** Definitions and counts Download ↓](https://nosible.com/data/world-v1.2/ontologies/iptc-genre/v1/nodes.csv) [**JSON** Full machine-readable release Download ↓](https://nosible.com/data/world-v1.2/ontologies/iptc-genre/v1/statistics.json) Release details and citation files [Manifest](https://nosible.com/data/world-v1.2/ontologies/iptc-genre/v1/manifest.json) [Release README](https://nosible.com/data/world-v1.2/ontologies/iptc-genre/v1/README.md) [CC0 license and scope](https://nosible.com/data/world-v1.2/ontologies/iptc-genre/v1/LICENSE.md) [Citation file](https://nosible.com/data/world-v1.2/ontologies/iptc-genre/v1/CITATION.cff) [Changelog](https://nosible.com/data/world-v1.2/ontologies/iptc-genre/v1/CHANGELOG.md) Need the complete World event schema? [Open the World data dictionary.](https://nosible.com/data-dictionaries#world) ### References 1. [ 1 ][International Press Telecommunications Council. NewsCodes Genre controlled categories.](https://www.iptc.org/std/NewsCodes/treeview/genre/genre-en-GB.html) 2. [ 2 ][Board of Governors of the Federal Reserve System. Federal Open Market Committee meeting calendars and information.](https://www.federalreserve.gov/monetarypolicy/fomccalendars.htm) 3. [ 3 ][Lucca, D. O., & Moench, E. (2015). The pre-FOMC announcement drift. Journal of Finance, 70(1), 329–371.](https://doi.org/10.1111/jofi.12196) 4. [ 4 ][Nakamura, E., & Steinsson, J. (2018). High-frequency identification of monetary non-neutrality: The information effect. Quarterly Journal of Economics, 133(3), 1283–1330.](https://doi.org/10.1093/qje/qjy004) 5. [ 5 ][Binder, C. (2018). Federal Reserve communication and the media. Journal of Media Economics, 30(4), 191–214.](https://doi.org/10.1080/08997764.2018.1515767) 6. [ 6 ][Banerjee, S., Cordova, P., De Pooter, M., & Grishchenko, O. V. (2025). Gauging the sentiment of Federal Open Market Committee communications through the eyes of the financial press. FEDS 2025-048.](https://doi.org/10.17016/FEDS.2025.048) 7. [ 7 ][UK Government. (2022). Prime Minister's statement on the death of Her Majesty Queen Elizabeth II.](https://www.gov.uk/government/speeches/prime-ministers-statement-on-the-death-of-her-majesty-queen-elizabeth-ii) 8. [ 8 ][Apple. (2024). Apple introduces iPhone 16 and iPhone 16 Plus.](https://www.apple.com/newsroom/2024/09/apple-introduces-iphone-16-and-iphone-16-plus/) Continue exploring ## Complementary ontologies [26 categories Schema.org Event Types Combine editorial format with a standard event clock for scheduled conferences, releases and public events. Explore ontology →](https://nosible.com/ontologies/schema-org-events) [718 categories IPTC Media Topics Separate how a story is written from the subject the story actually covers. Explore ontology →](https://nosible.com/ontologies/iptc-media-topics) > Explore 55 IPTC News Genre categories, World V1.2 evidence and example usages showing how forecasts, obituaries and supplied formats change event research. **URL:** https://nosible.com/ontologies/iptc-genre --- --- title: "IPTC Media Topics" description: "Explore 718 IPTC Media Topic categories and example usages showing how COVID-19, Suez and GameStop reporting migrate across distinct subject regimes." url: "https://nosible.com/ontologies/iptc-media-topics" --- [Home](https://nosible.com/) [Ontologies](https://nosible.com/ontologies) IPTC Media Topics Ontology field guide By NOSIBLE Research Updated 2026-07-18 # IPTC Media Topics IPTC Media Topics reveals when issuer coverage crosses risk domains. Subject labels determine the risk regime a researcher sees. An issuer can move from routine company and engineering coverage into accidents, regulation, litigation and workforce disruption. IPTC Media Topics preserves those domains in one hierarchy. Researchers can detect the transition without treating every company event as equivalent across time, sources, sectors or risk regimes. [[ 1 ]](https://www.iptc.org/std/NewsCodes/treeview/mediatopic/mediatopic-en-GB.html) World assigns each event a three-level topic path. Roots support broad regime measures. Leaves support precise cohorts. Topic transitions show where reporting moves. They do not establish economic impact, causal exposure, operational loss, legal liability or subsequent market repricing. [[ 1 ]](https://www.iptc.org/std/NewsCodes/treeview/mediatopic/mediatopic-en-GB.html) Categories 718 categories Structure 3 levels World events labelled 100.0% Stable codes Since v1 Release data CC0 On this page [01 Foundations](https://nosible.com/ontologies/iptc-media-topics#foundations) [02 Categories](https://nosible.com/ontologies/iptc-media-topics#vocabulary) [03 Example usages](https://nosible.com/ontologies/iptc-media-topics#trends) [04 Downloads and references](https://nosible.com/ontologies/iptc-media-topics#downloads) Foundations ## Media Topics separates subject domains without measuring importance, impact or factual accuracy IPTC Media Topics is a subject hierarchy for news. Seventeen roots cover economy, politics, health, science, labour, society, environment, sport and other domains. Descendants add narrower subjects. One structure supports broad attention measures and precise event cohorts for direct comparison. [[ 1 ]](https://www.iptc.org/std/NewsCodes/treeview/mediatopic/mediatopic-en-GB.html) World stores the dominant path for each canonical event. It does not preserve every secondary subject in the underlying coverage. A broader topic mix can reflect real diffusion, changing sources or classification boundaries. Validate important transitions against entities, sources and leaf assignments before interpreting the regime shift. [[ 1 ]](https://www.iptc.org/std/NewsCodes/treeview/mediatopic/mediatopic-en-GB.html) Categories ## Media Topics organizes 718 reporting categories beneath 17 independent subject-domain roots World implements 17 roots, 73 intermediate subjects and 628 leaves. Each event receives one path. The explorer exposes definitions, stable IPTC codes, complete paths and World V1.2 frequencies. Bounded search and tree expansion avoid rendering all 718 categories at once while preserving the full hierarchy. Search the ontology 17 / 718 shown World V1.2 coverage Coverage **100.0%** Labelled events **15,311,040** World events **15,311,040** Ontology index 15,311,040 of 15,311,040 World events carry IPTC Media Topics labels. Select a category to inspect the evidence Hierarchy level_1 / level_2 / level_3 › arts, culture, entertainment and media Broad media, culture and entertainment attention · 3 children › conflict, war and peace medtop:16000000 · 9 children › crime, law and justice medtop:02000000 · 5 children › disaster, accident and emergency incident medtop:03000000 · 5 children › economy, business and finance Commercial, industry and financial transmission · 5 children › education medtop:05000000 · 13 children › environment medtop:06000000 · 6 children › health Clinical, healthcare and population-health cohorts · 10 children › human interest medtop:08000000 · 10 children › labour Employment, compensation and workforce effects · 7 children › lifestyle and leisure medtop:10000000 · 3 children › politics and government Regulation, public spending and policy response · 8 children › religion medtop:12000000 · 10 children › science and technology medtop:13000000 · 8 children › society Social adoption, access and distributional effects · 13 children › sport Competition, scheduling and sports-business effects · 12 children › weather medtop:17000000 · 4 children arts, culture, entertainment and media ### arts, culture, entertainment and media Copy link ↗ #### Definition Reporting about arts, culture, entertainment and media, including creative work, cultural practice, publishing and broadcasting. Research use Broad media, culture and entertainment attention Code medtop:01000000 Events with label 770,168 Share of labelled 5.0% Share of World 5.0% **Events with label** is the selected count. **Label share** divides it by 15,311,040 assigned events; **World share** divides it by all 15,311,040 events. #### Representative World V1.2 events One strong classified example per available year, with up to ten years shown. 10 examples 1. 2015-12-04Local Arts Groups Propose Downtown Cultural District RevitalizationCoverage 11 2. 2016-01-13Corus Acquires Shaw Media for $2.65B to Expand Canadian TV PortfolioCoverage 73 3. 2017-01-19Social Media Cancels Hollywood Events Including Performers and ShowsCoverage 39 4. 2019-02-20Ubisoft Developing Female-Led Skull Bones TV Series with Atlas EntertainmentCoverage 35 5. 2020-07-28Sky Arts Becomes Free-to-Air on Freeview and Freesat in SeptemberCoverage 37 6. 2021-10-08UK City of Culture 2025 Longlist Revealed for Eight AreasCoverage 54 7. 2022-09-19Naples Launches ¡Arte Viva! Hispanic Culture Arts InitiativeCoverage 52 8. 2024-01-17Social Media Algorithms Flatten Culture by Making DecisionsCoverage 234 9. 2025-06-17ARKO Concludes 10th World Summit on Arts and Culture in SeoulCoverage 68 10. 2026-04-30Teens Prefer Social Media and Influencers for News Despite SkepticismCoverage 366 Browse categories to compare definitions, World V1.2 statistics and representative classified events. Example usage 1 of 3 ### Example Usage: COVID-19 coverage escapes health as policy, society and sport take over COVID-19 reporting begins as a health story, then spreads across the economy and public life. Health falls from 37.4% of assigned topics before 20 February 2020 to 23.0% afterwards. Politics, society and sport each gain material share as governments restrict activity, close borders and cancel events. [[ 2 ]](https://www.who.int/news/item/29-06-2020-covidtimeline) Each topic implies a different exposure set. Health points to providers and therapeutics. Politics identifies regulation and fiscal response. Society captures behavioral disruption. Sport isolates shutdowns and media rights. The hierarchy turns one pandemic keyword into separate event clocks for sector, policy and cross-asset research. [[ 1 ]](https://www.iptc.org/std/NewsCodes/treeview/mediatopic/mediatopic-en-GB.html) [[ 3 ]](https://www.imf.org/en/Publications/WEO/Issues/2020/04/14/weo-april-2020) Health **37.4% to 23.0%** Politics **+5.8pp** Society **+5.6pp** Sport **+4.0pp** **COVID-19 expands from a health story into policy, society and sport** two-week share of assigned topics Health Politics Society Sport The fixed English cohort contains 55,650 IPTC-assigned COVID-19 events from January through June 2020. Lines pool topic counts across consecutive weeks and divide by all assigned cohort events. They measure the reporting subject mix, not infections, economic losses or market returns. Example usage 2 of 3 ### Example Usage: Suez topics separate the immediate trade shock from the longer logistics bottleneck The Ever Given blockage creates a concentrated trade and transport regime after 23 March 2021. Normalized IPTC-assigned coverage reaches 2,800 events per million during the blocked week. Economy, Business and Finance accounts for 64.3% of assigned topics, Transport for 35.7%, and International Trade for 14.3% while the canal remains closed. [[ 6 ]](https://www.suezcanal.gov.eg/English/MediaCenter/News/Pages/sca_29-3-2021.aspx) After refloating, normalized attention falls to 584 events per million. Economy and Finance drops to 41.0%, while Transport remains at 38.5% and infrastructure-related Science and Technology reaches 7.7%. The categories create separate clocks for the initial trade interruption and the slower freight, inventory, insurance and port-congestion consequences. [[ 7 ]](https://unctad.org/news/suez-canal-blockage-hit-global-trade-6-billion-9-billion-week) Normalized coverage -2,216 per million Economy and finance -23.3pp Transport +2.8pp International trade -14.3pp Three-week pooled share of IPTC-assigned Suez events · % Economy and Finance Transport Science and Technology The fixed English cohort requires Ever Given or Suez Canal language from March through mid-May 2021. Each line pools category counts across the current and prior two weeks, then divides by all IPTC-assigned cohort events in that window. Economy and Science are top-level topics; Transport is the relevant third-level topic within the same fixed canal-disruption reporting cohort throughout the event window. Example usage 3 of 3 ### Example Usage: GameStop becomes a policy story when Congress examines the trading restrictions GameStop begins as a market and retail-investor story. Economy, Business and Finance represents 41% of assigned topics around the 28 January 2021 trading restrictions. Its share falls to 29% by the 18 February congressional hearing, while Politics and Government rises from 5% to 21% of the same classified cohort. [[ 8 ]](https://www.sec.gov/files/staff-report-equity-options-market-struction-conditions-early-2021.pdf) The shift separates two risk channels that a GameStop keyword combines. The squeeze concerns positioning, liquidity and price formation. The hearing introduces brokerage rules, payment for order flow and market-structure risk. Topic labels identify when the event clock moves from trading mechanics toward regulation and business-model risk. [[ 9 ]](https://financialservices.house.gov/calendar/eventsingle.aspx?EventID=407107) Politics +16.0pp Economy and Finance -12.0pp Lifestyle and Leisure +11.0pp Fourteen-day pooled share of IPTC-assigned GameStop events · % Economy and Finance Politics and Government Lifestyle and Leisure The fixed English cohort covers January through April 2021 and requires GameStop, Game Stop or GME. The chart shows 15 January through 10 March. Each line pools topic counts across the current and prior thirteen days, then divides by all IPTC-assigned cohort events. The cohort contains 231 assigned events; raw counts are sample disclosures only and do not directly establish market impact or causal effects. Data and sources ## Download IPTC Media Topics and References ### Downloads Download categories and counts as CSV, or the complete machine-readable release as JSON. [**CSV** Definitions and counts Download ↓](https://nosible.com/data/world-v1.2/ontologies/iptc-media-topics/v1/nodes.csv) [**JSON** Full machine-readable release Download ↓](https://nosible.com/data/world-v1.2/ontologies/iptc-media-topics/v1/statistics.json) Release details and citation files [Manifest](https://nosible.com/data/world-v1.2/ontologies/iptc-media-topics/v1/manifest.json) [Release README](https://nosible.com/data/world-v1.2/ontologies/iptc-media-topics/v1/README.md) [CC0 license and scope](https://nosible.com/data/world-v1.2/ontologies/iptc-media-topics/v1/LICENSE.md) [Citation file](https://nosible.com/data/world-v1.2/ontologies/iptc-media-topics/v1/CITATION.cff) [Changelog](https://nosible.com/data/world-v1.2/ontologies/iptc-media-topics/v1/CHANGELOG.md) Need the complete World event schema? [Open the World data dictionary.](https://nosible.com/data-dictionaries#world) ### References 1. [ 1 ][International Press Telecommunications Council. Media Topics controlled categories, 2025-10-10 release.](https://www.iptc.org/std/NewsCodes/treeview/mediatopic/mediatopic-en-GB.html) 2. [ 2 ][World Health Organization. (2020). Timeline of WHO's response to COVID-19.](https://www.who.int/news/item/29-06-2020-covidtimeline) 3. [ 3 ][International Monetary Fund. (2020). World Economic Outlook: The Great Lockdown.](https://www.imf.org/en/Publications/WEO/Issues/2020/04/14/weo-april-2020) 4. [ 8 ][U.S. Securities and Exchange Commission. (2021). Staff report on equity and options market structure conditions in early 2021.](https://www.sec.gov/files/staff-report-equity-options-market-struction-conditions-early-2021.pdf) 5. [ 9 ][U.S. House Committee on Financial Services. (2021). Game Stopped? Who wins and loses when short sellers, social media, and retail investors collide.](https://financialservices.house.gov/calendar/eventsingle.aspx?EventID=407107) 6. [ 6 ][Suez Canal Authority. (2021). Successful refloating of the Ever Given, 29 March 2021.](https://www.suezcanal.gov.eg/English/MediaCenter/News/Pages/sca_29-3-2021.aspx) 7. [ 7 ][United Nations Conference on Trade and Development. (2021). Suez Canal blockage could affect global trade by $6-9 billion per week.](https://unctad.org/news/suez-canal-blockage-hit-global-trade-6-billion-9-billion-week) Continue exploring ## Complementary ontologies [709 categories IAB Content Ontology Cross-check news topics against a separate web-content hierarchy for broader subject coverage. Explore ontology →](https://nosible.com/ontologies/iab-content-taxonomy) [5 categories Media Frames Measure how a subject is framed after identifying what the reporting is about. Explore ontology →](https://nosible.com/ontologies/media-frames) > Explore 718 IPTC Media Topic categories and example usages showing how COVID-19, Suez and GameStop reporting migrate across distinct subject regimes. **URL:** https://nosible.com/ontologies/iptc-media-topics --- --- title: "Media Frames" description: "Explore five media-frame categories and evidence showing how conflict, responsibility, economic consequence and human-interest reporting separate event risks." url: "https://nosible.com/ontologies/media-frames" --- [Home](https://nosible.com/) [Ontologies](https://nosible.com/ontologies) Media Frames Ontology field guide By NOSIBLE Research Updated 2026-07-18 # Media Frames Media Frames separates an event from the narrative used to explain it. The same event can be reported as a financial consequence, conflict, personal story, moral question or assignment of responsibility. Those choices create different variables even when the company, date and event stay fixed. Media Frames records the dominant narrative for direct comparison. [[ 1 ]](https://doi.org/10.1111/j.1460-2466.2000.tb02843.x) [[ 2 ]](https://doi.org/10.1111/j.0021-9916.2007.00326.x) World assigns one frame to every canonical event. Researchers can segment event regimes and test whether framing adds information beyond topic, sentiment and event type. This guide provides definitions, World base rates, a Wirecard study and downloadable observations for independent analysis. Categories 5 categories Structure Flat, no hierarchy World events labelled 100.0% Stable codes Since v1 Release data CC0 On this page [01 Foundations](https://nosible.com/ontologies/media-frames#foundations) [02 Categories](https://nosible.com/ontologies/media-frames#vocabulary) [03 Example usages](https://nosible.com/ontologies/media-frames#trends) [04 Downloads and references](https://nosible.com/ontologies/media-frames#downloads) Foundations ## Media Frames separates narrative emphasis from topic, sentiment and factual accuracy Economic Consequences foregrounds costs and market effects. Responsibility assigns cause or duty. Conflict foregrounds disagreement. Human Interest centres people. Morality evaluates conduct through ethical or religious principles. The labels describe presentation rather than subject. This distinction is directly testable. [[ 1 ]](https://doi.org/10.1111/j.1460-2466.2000.tb02843.x) World classifies the dominant frame in each event record. It does not infer editorial intent or prove market impact. One label compresses mixed narratives. Test disagreement near class boundaries, source concentration and alternative text representations before using the field in a model across time and source regimes. [[ 2 ]](https://doi.org/10.1111/j.0021-9916.2007.00326.x) Categories ## Five peer labels separate market effects, conflict, responsibility, people and morality The ontology is flat. Every World event receives one of five labels. The explorer provides each definition, stable code and exact World V1.2 count. Economic Consequences is the largest class with 4,760,325 events. Morality is the smallest with 748,524. Flat labels keep comparisons direct across events, sources and time. [[ 1 ]](https://doi.org/10.1111/j.1460-2466.2000.tb02843.x) Search the ontology 5 / 5 shown World V1.2 coverage Coverage **100.0%** Labelled events **15,311,040** World events **15,311,040** Ontology index 15,311,040 of 15,311,040 World events carry Media Frames labels. Select a category to inspect the evidence Categories media_frame Conflict Institutional disagreement, competing interests, strategic opposition and confrontation between parties Human Interest Individual experience and personal consequences Economic Consequences Financial costs, market effects and resource allocation Morality Ethical, religious or values-based evaluation Responsibility Attributed cause, obligation and capacity to resolve a problem Conflict ### Conflict Copy link ↗ #### Definition Reporting organized around disagreement or incompatible goals between political, commercial and institutional parties. Research use Institutional disagreement, competing interests, strategic opposition and confrontation between parties Code Conflict Events with label 4,183,358 Share of labelled 27.3% Share of World 27.3% **Events with label** is the selected count. **Label share** divides it by 15,311,040 assigned events; **World share** divides it by all 15,311,040 events. #### Representative World V1.2 events One strong classified example per available year, with up to ten years shown. 10 examples 1. 2015-07-10BlueIndy Charging Stations Ignite Conflict Between Mayor and CouncilmanCoverage 10 2. 2016-01-28Arizona Regulator Burns Investigates APS Political Spending Amid ConflictCoverage 28 3. 2017-11-07Neighbor Charged With Assault After Conflict With US Senator Rand PaulCoverage 106 4. 2019-05-13EU and UK Warn US-Iran Conflict Could Erupt by AccidentCoverage 88 5. 2020-03-02Turkey Syria Conflict Drives Migrant Crisis at Greek BorderCoverage 192 6. 2021-05-20Israel and Hamas Agree Ceasefire Ending 11-Day Gaza ConflictCoverage 1,400 7. 2022-11-30Schools in Purple Districts Face Rising Political Conflict and HateCoverage 133 8. 2024-09-24Israeli Strikes Kill Nearly 500 in Lebanon Amid Escalating Hezbollah ConflictCoverage 3,381 9. 2025-05-16Trump Plans Direct Talks with Putin to End Ukraine ConflictCoverage 3,087 10. 2026-03-02Middle East War Drives Gas Prices Up as Iran Conflict WidensCoverage 60,216 Browse categories to compare definitions, World V1.2 statistics and representative classified events. Example usage 1 of 3 ### Example Usage: Wirecard reporting moves from dispute before collapse to public accountability afterward The cohort contains 170 English World events that explicitly name Wirecard from January 2019 through September 2020. The collapse window begins on 18 June 2020, when the company disclosed that €1.9 billion in cash could not be confirmed. Wirecard filed for insolvency on 25 June, one week after the disclosure. [[ 3 ]](https://www.bafin.de/SharedDocs/Downloads/EN/Jahresbericht/dl_jb_2020_en.pdf?__blob=publicationFile&v=2) [[ 5 ]](https://www.bafin.de/SharedDocs/Veroeffentlichungen/EN/Pressemitteilung/2020/pm_200625_wirecard_en.html) Conflict accounts for 36.8% of 95 pre-collapse events and falls to 9.3% across 75 collapse events. Responsibility rises from 5.3% to 32.0%, while Economic Consequences remains the largest frame. The transition separates prolonged dispute from assigned accountability after the balance-sheet failure became public, which polarity cannot identify. [[ 1 ]](https://doi.org/10.1111/j.1460-2466.2000.tb02843.x) [[ 4 ]](https://www.esma.europa.eu/press-news/esma-news/esma-identifies-deficiencies-in-german-supervision-wirecards-financial-reporting) Explicit Wirecard cohort · 30-day memory 170 events · 2019-01-01 to 2020-09-30 **Responsibility replaces conflict after Wirecard discloses missing cash** share of recent classified events Conflict Responsibility Economic Consequences Each line is the exponentially weighted share of recent Wirecard events assigned to that frame. The thirty-day half-life suppresses isolated daily observations while preserving the transition around the missing-cash disclosure. The chart describes reporting emphasis, not fraud timing, causality or market impact. Example usage 2 of 3 ### Example Usage: Layoff coverage adds blame and conflict that earnings data cannot explain Earnings beats are almost entirely framed through Economic Consequences. Workforce reductions are not. Responsibility, Conflict and Human Interest rise from a combined 1.4% of earnings-beat events to 12.7% of layoff events. The underlying cost action therefore carries information that reported earnings alone cannot measure. [[ 1 ]](https://doi.org/10.1111/j.1460-2466.2000.tb02843.x) That distinction separates expected margin improvement from execution and reputation risk. A portfolio researcher can test whether blame, conflict or employee-centred coverage predicts customer loss, slower hiring, litigation or weaker revisions after announced cuts. Media Frames adds that narrative channel without replacing the underlying event classification. [[ 2 ]](https://doi.org/10.1111/j.0021-9916.2007.00326.x) Responsibility +4.3pp Conflict +4.8pp Human Interest +2.3pp Economic Consequences -12.1pp Composition of classified events share of each cohort **Earnings beats** 98.6 % **Workforce reductions** 13.5 % 86.5 % Non-economic framing Economic consequences The fixed Q1 2024 English panel contains 919 earnings-beat events and 561 workforce-reduction events. Counts are divided by all eligible English World events; frame values are within-class shares for direct Q1 comparison between the cohorts. Example usage 3 of 3 ### Example Usage: Responsibility framing identifies the liability clock after a disclosed data breach Earnings misses and data breaches are both negative corporate events, but their frames differ sharply. Economic Consequences accounts for 99.7% of missed-earnings events. Responsibility reaches 45.2% of breach disclosures and Human Interest reaches 20.2%. Polarity cannot separate financial disappointment from breach-related blame inside one sentiment score, leaving distinct liability channels pooled after public disclosure. [[ 1 ]](https://doi.org/10.1111/j.1460-2466.2000.tb02843.x) The frame mix creates a liability clock. Responsibility and Human Interest can be tested against regulatory action, litigation, customer attrition, insurance claims and persistent attention after the initial disclosure. Media Frames isolates those channels while NOSIBLE Events holds the event class constant, reducing generic negative-news contamination. [[ 2 ]](https://doi.org/10.1111/j.0021-9916.2007.00326.x) Responsibility +45.2pp Human Interest +20.2pp Economic Consequences -77.1pp Conflict +9.4pp Composition of classified events share of each cohort **Earnings misses** 99.7 % **Disclosed breaches** 45.2 % 22.6 % 20.2 % 12.0 % Responsibility Economic consequences Human interest Conflict and other The fixed Q1 2024 English panel contains 327 earnings-miss events and 124 disclosed-breach events. Media Frames supplies within-class narrative shares; NOSIBLE Events supplies the corporate event classes within the same quarterly panel. Data and sources ## Download Media Frames and References ### Downloads Download categories and counts as CSV, or the complete machine-readable release as JSON. [**CSV** Definitions and counts Download ↓](https://nosible.com/data/world-v1.2/ontologies/media-frames/v1/nodes.csv) [**JSON** Full machine-readable release Download ↓](https://nosible.com/data/world-v1.2/ontologies/media-frames/v1/statistics.json) Release details and citation files [Manifest](https://nosible.com/data/world-v1.2/ontologies/media-frames/v1/manifest.json) [Release README](https://nosible.com/data/world-v1.2/ontologies/media-frames/v1/README.md) [CC0 license and scope](https://nosible.com/data/world-v1.2/ontologies/media-frames/v1/LICENSE.md) [Citation file](https://nosible.com/data/world-v1.2/ontologies/media-frames/v1/CITATION.cff) [Changelog](https://nosible.com/data/world-v1.2/ontologies/media-frames/v1/CHANGELOG.md) Need the complete World event schema? [Open the World data dictionary.](https://nosible.com/data-dictionaries#world) ### References 1. [ 1 ][Semetko, H. A., & Valkenburg, P. M. (2000). Framing European politics: A content analysis of press and television news. Journal of Communication, 50(2), 93-109.](https://doi.org/10.1111/j.1460-2466.2000.tb02843.x) 2. [ 2 ][Scheufele, D. A., & Tewksbury, D. (2007). Framing, agenda setting, and priming. Journal of Communication, 57(1), 9-20.](https://doi.org/10.1111/j.0021-9916.2007.00326.x) 3. [ 3 ][Federal Financial Supervisory Authority. (2021). Annual Report 2020, Wirecard chronology and supervisory review.](https://www.bafin.de/SharedDocs/Downloads/EN/Jahresbericht/dl_jb_2020_en.pdf?__blob=publicationFile&v=2) 4. [ 4 ][European Securities and Markets Authority. (2020). ESMA identifies deficiencies in German supervision of Wirecard's financial reporting.](https://www.esma.europa.eu/press-news/esma-news/esma-identifies-deficiencies-in-german-supervision-wirecards-financial-reporting) 5. [ 5 ][Federal Financial Supervisory Authority. (2020). Statement on Wirecard and the company's insolvency filing.](https://www.bafin.de/SharedDocs/Veroeffentlichungen/EN/Pressemitteilung/2020/pm_200625_wirecard_en.html) Continue exploring ## Complementary ontologies [7 categories Ekman-7 Emotion Ontology Compare narrative emphasis with the distinct emotional frame expressed in the same event reporting. Explore ontology →](https://nosible.com/ontologies/ekman-emotion) [718 categories IPTC Media Topics Separate a story's narrative frame from its subject to avoid confusing presentation with topic. Explore ontology →](https://nosible.com/ontologies/iptc-media-topics) > Explore five media-frame categories and evidence showing how conflict, responsibility, economic consequence and human-interest reporting separate event risks. **URL:** https://nosible.com/ontologies/media-frames --- --- title: "MITRE ATT&CK Enterprise" description: "Explore 866 MITRE ATT&CK categories and example usages separating ransomware persistence, exfiltration and operational impact across major cyber incidents." url: "https://nosible.com/ontologies/mitre-attack" --- [Home](https://nosible.com/) [Ontologies](https://nosible.com/ontologies) MITRE ATT&CK Ontology field guide By NOSIBLE Research Updated 2026-07-18 # MITRE ATT&CK Enterprise MITRE ATT&CK separates reported cyber operations by objective, technique and implementation detail across an attack sequence. Cyber incidents are not interchangeable. Credential theft, persistence, lateral movement, command and control, exfiltration and operational impact describe different parts of an intrusion. ATT&CK provides a common hierarchy. Tactics state why an adversary acts. Techniques state how. Sub-techniques add implementation detail. [[ 1 ]](https://attack.mitre.org/resources/) [[ 2 ]](https://attack.mitre.org/docs/ATTACK_Design_and_Philosophy_March_2020.pdf) World applies Enterprise ATT&CK to event reporting. Researchers can test whether technique mix improves incident triage, vendor-risk estimates or market-response models beyond a generic cyber-event flag. This guide provides the hierarchy, World coverage, three example usages and downloadable observations. [[ 4 ]](https://www.sec.gov/rules-regulations/2023/07/s7-09-22) [[ 5 ]](https://doi.org/10.1016/j.jfineco.2019.05.019) [[ 6 ]](https://doi.org/10.3386/w28906) Categories 866 categories Structure 3 levels World events labelled 9.7% Stable codes Since v1 Release data CC0 On this page [01 Foundations](https://nosible.com/ontologies/mitre-attack#foundations) [02 Categories](https://nosible.com/ontologies/mitre-attack#vocabulary) [03 Example usages](https://nosible.com/ontologies/mitre-attack#trends) [04 Downloads and references](https://nosible.com/ontologies/mitre-attack#downloads) Foundations ## ATT&CK classifies observed behaviour without measuring severity or attacker capability Enterprise ATT&CK organises observed adversary behaviour into tactics, techniques and sub-techniques. A tactic is an operational goal. A technique is a method used to reach that goal. A sub-technique is a more specific method. ATT&CK is maintained from documented real-world observations across known operations. [[ 1 ]](https://attack.mitre.org/resources/) [[ 2 ]](https://attack.mitre.org/docs/ATTACK_Design_and_Philosophy_March_2020.pdf) World classifies the behaviour described in an event record. It does not verify an attack, reconstruct the complete attack chain or infer attacker skill. A deeper path can mean that reporting contains more implementation detail. It does not prove that the attacker was more capable or the attack more damaging. [[ 1 ]](https://attack.mitre.org/resources/) Categories ## ATT&CK maps 866 attack categories across tactics, techniques and sub-techniques World uses Enterprise ATT&CK v16.1 with 14 tactic roots, 866 path-specific categories, 731 leaves and a maximum depth of two. Techniques can appear under several tactics because one method can serve several goals. The explorer preserves each path, definition and World count for direct comparison across phases. [[ 2 ]](https://attack.mitre.org/docs/ATTACK_Design_and_Philosophy_March_2020.pdf) Search the ontology 14 / 866 shown World V1.2 coverage Coverage **9.7%** Labelled events **1,488,001** World events **15,311,040** Ontology index 1,488,001 of 15,311,040 World events carry MITRE ATT&CK Enterprise labels. Select a category to inspect the evidence Hierarchy tactic / technique / subtechnique › Execution Code or command execution on a target system · 14 children › Collection Data gathered before theft or use · 17 children › Persistence Methods used to retain access · 20 children › Privilege Escalation Methods used to gain higher permissions · 14 children › Credential Access Passwords, tokens and authentication material · 17 children › Discovery Mapping systems and network structure · 32 children › Resource Development Infrastructure, accounts and capabilities prepared before access · 8 children › Reconnaissance Target selection and pre-operation information gathering · 10 children › Defense Evasion Methods used to avoid detection · 44 children › Initial Access Methods used to enter a target environment · 10 children › Impact Disruption, destruction or financial theft · 14 children › Lateral Movement Movement between systems · 9 children › Command and Control Communication with compromised systems · 18 children › Exfiltration Data removed from the target · 9 children Execution ### Execution Copy link ↗ #### Definition Adversaries execute code or commands on target systems. Research use Code or command execution on a target system Code Execution Events with label 46,109 Share of labelled 3.1% Share of World 0.3% **Events with label** is the selected count. **Label share** divides it by 1,488,001 assigned events; **World share** divides it by all 15,311,040 events. #### Representative World V1.2 events One strong classified example per available year, with up to ten years shown. 10 examples 1. 2015-07-07HSBC Fires Six Staff Over Mock ISIS Execution Video in BirminghamCoverage 88 2. 2016-05-20Oklahoma Grand Jury Condemns Careless Execution Drug Mix-UpCoverage 32 3. 2017-02-28Critical ESET Antivirus Flaw Enables Mac Remote Code ExecutionCoverage 16 4. 2019-06-22Iran Executes Former Defense Contractor for Alleged CIA SpyingCoverage 39 5. 2020-07-12Federal Execution Prep Worker Tests Positive for CoronavirusCoverage 19 6. 2021-02-02Alabama Faces COVID-19 Execution Risks for SmithCoverage 53 7. 2022-11-16Death Penalty Execution Workers Face Secret Toll and Political ShiftsCoverage 317 8. 2024-10-28Iran Executes German-Iranian Dissident Jamshid Sharmahd on Terrorism ChargesCoverage 159 9. 2025-05-08South Carolina Botched Firing Squad Execution Causes Extreme PainCoverage 89 10. 2026-06-25British Influencer Faces Execution in Dubai for Alleged Murder of PartnerCoverage 96 Browse categories to compare definitions, World V1.2 statistics and representative classified events. Example usage 1 of 3 ### Example Usage: Three cyber incidents clearly expose different operational and financial risk channels SolarWinds, MOVEit and Change Healthcare are major cyber incidents, but their ATT&CK profiles differ. Entry preparation accounts for 60% of SolarWinds tactic assignments. Exfiltration reaches 20% for MOVEit. Impact accounts for 63% of Change Healthcare assignments after the ransomware disruption affected payment and claims infrastructure. [[ 3 ]](https://www.cisa.gov/news-events/cybersecurity-advisories/aa20-352a) [[ 7 ]](https://www.cisa.gov/news-events/cybersecurity-advisories/aa23-158a) [[ 8 ]](https://www.unitedhealthgroup.com/ns/changehealthcare.html) These three incidents create different exposure maps. SolarWinds shows software and credential risk. MOVEit shows stolen-data liability. Change Healthcare shows operational cash flow and service disruption. ATT&CK separates these mechanisms before tests of suppliers, customers, insurers, spreads, revisions and incident duration across firms and time. [[ 1 ]](https://attack.mitre.org/resources/) [[ 5 ]](https://doi.org/10.1016/j.jfineco.2019.05.019) [[ 6 ]](https://doi.org/10.3386/w28906) SolarWinds **60% entry preparation** MOVEit **20% exfiltration** Change Healthcare **63% impact** **Three cyber incidents expose different operational risk channels** share of ATT&CK tactic assignments **SolarWinds** third-party access · n= 114 60 % 38 % **MOVEit** data theft · n= 55 13 % 27 % 20 % 27 % 13 % **Change Healthcare** operational disruption · n= 73 10 % 63 % 27 % Entry preparation Command and control Exfiltration Impact Other tactics Each fixed English cohort requires an explicit incident name. The chart pools all ATT&CK tactic assignments inside the stated study windows. Entry preparation combines Reconnaissance, Resource Development and Initial Access. It describes classified reporting, not confirmed attack paths, losses or attacker capability. Example usage 2 of 3 ### Example Usage: Ransomware impact fades quickly in 2017 but persists after Change Healthcare WannaCry and NotPetya create sharp Impact-classified reporting spikes in 2017. Their seven-day normalized rates peak at 1,644 and 883 events per million, then return to zero within 30 days. Change Healthcare peaks later at 536 per million and remains elevated 60 days after the February 2024 disruption. [[ 9 ]](https://www.cisa.gov/news-events/alerts/2017/05/12/indicators-associated-wannacry-ransomware) [[ 10 ]](https://www.cisa.gov/news-events/cybersecurity-advisories/aa17-181a) Peak attention alone misses the difference. A fast global malware wave and a persistent payment-system outage create different cash-flow, counterparty and operational-risk windows. ATT&CK's Impact tactic provides a common behavioural denominator for comparing incident duration without raw coverage or generic cyber keywords. [[ 11 ]](https://www.unitedhealthgroup.com/ns/changehealthcare.html) Change day 60 +143 per million WannaCry peak +1,644 per million Seven-day normalized ATT&CK Impact rate by event day · per million Change Healthcare WannaCry NotPetya Each line aligns a fixed English incident cohort to public disclosure day. Counts assigned to the ATT&CK Impact tactic are pooled across the current and prior six days, divided by all English World events in the same window and plotted from event day -14 to +60. Example usage 3 of 3 ### Example Usage: MOVEit reporting shifts from command traffic into exfiltration and operational impact MOVEit reporting changes after the initial vulnerability disclosure. Command and Control dominates the first month. Exfiltration rises from 4.8% to 23.8% during the follow-on period, while Transfer Data to Cloud Account reaches 19.1%. The classified focus moves from exploit communication toward stolen-data handling and transfer paths. [[ 7 ]](https://www.cisa.gov/sites/default/files/2023-06/aa23-158a-stopransomware-cl0p-ransomware-gang-exploits-moveit-vulnerability_2.pdf) That shift changes the financial question. Vulnerability disclosure concerns immediate exposure and remediation. Exfiltration creates notification, litigation, customer and insurance liabilities that can surface later. ATT&CK supplies the transition clock, starting exposed-company tests when reported behaviour changes rather than at the initial breach headline or disclosure before data theft appears across exposed firms. [[ 8 ]](https://www.cisa.gov/known-exploited-vulnerabilities-catalog) Exfiltration +19.0pp Transfer to cloud +14.3pp Global baseline +0.2pp Normalized coverage -243 per million Fifty-six-day pooled share of ATT&CK tactic assignments · % Exfiltration Command and Control Impact The fixed English MOVEit cohort contains 21 ATT&CK-assigned events from 31 May to 30 June 2023 and 34 from July through December. The chart shows 15 June through 31 August. Lines pool tactic counts across the current and prior 55 days, then divide by all tactic assignments in the same window. Data and sources ## Download MITRE ATT&CK and References ### Downloads Download categories and counts as CSV, or the complete machine-readable release as JSON. [**CSV** Definitions and counts Download ↓](https://nosible.com/data/world-v1.2/ontologies/mitre-attack/v1/nodes.csv) [**JSON** Full machine-readable release Download ↓](https://nosible.com/data/world-v1.2/ontologies/mitre-attack/v1/statistics.json) Release details and citation files [Manifest](https://nosible.com/data/world-v1.2/ontologies/mitre-attack/v1/manifest.json) [Release README](https://nosible.com/data/world-v1.2/ontologies/mitre-attack/v1/README.md) [CC0 license and scope](https://nosible.com/data/world-v1.2/ontologies/mitre-attack/v1/LICENSE.md) [Citation file](https://nosible.com/data/world-v1.2/ontologies/mitre-attack/v1/CITATION.cff) [Changelog](https://nosible.com/data/world-v1.2/ontologies/mitre-attack/v1/CHANGELOG.md) Need the complete World event schema? [Open the World data dictionary.](https://nosible.com/data-dictionaries#world) ### References 1. [ 1 ][MITRE. ATT&CK: Get Started. Enterprise ATT&CK knowledge base and usage guidance.](https://attack.mitre.org/resources/) 2. [ 2 ][Strom, B. E. et al. (2020). The Design and Philosophy of MITRE ATT&CK. MITRE Technical Report MP180360R1.](https://attack.mitre.org/docs/ATTACK_Design_and_Philosophy_March_2020.pdf) 3. [ 3 ][Cybersecurity and Infrastructure Security Agency. (2020). Advanced Persistent Threat Compromise of Government Agencies, Critical Infrastructure, and Private Sector Organizations.](https://www.cisa.gov/news-events/cybersecurity-advisories/aa20-352a) 4. [ 4 ][U.S. Securities and Exchange Commission. (2023). Cybersecurity Risk Management, Strategy, Governance, and Incident Disclosure. Release 33-11216.](https://www.sec.gov/rules-regulations/2023/07/s7-09-22) 5. [ 5 ][Kamiya, S., Kang, J.-K., Kim, J., Milidonis, A., & Stulz, R. M. (2021). Risk management, firm reputation, and the impact of successful cyberattacks on target firms. Journal of Financial Economics, 139(3), 719-749.](https://doi.org/10.1016/j.jfineco.2019.05.019) 6. [ 6 ][Jamilov, R., Rey, H., & Tahoun, A. (2025). The Anatomy of Cyber Risk. NBER Working Paper 28906, revised December 2025.](https://doi.org/10.3386/w28906) 7. [ 7 ][Cybersecurity and Infrastructure Security Agency and Federal Bureau of Investigation. (2023). CL0P ransomware gang exploits MOVEit vulnerability.](https://www.cisa.gov/news-events/cybersecurity-advisories/aa23-158a) 8. [ 11 ][UnitedHealth Group. (2024). Change Healthcare cyber response and restoration updates.](https://www.unitedhealthgroup.com/ns/changehealthcare.html) 9. [ 9 ][Cybersecurity and Infrastructure Security Agency. (2017). Indicators associated with WannaCry ransomware.](https://www.cisa.gov/news-events/alerts/2017/05/12/indicators-associated-wannacry-ransomware) 10. [ 10 ][Cybersecurity and Infrastructure Security Agency. (2017). Petya ransomware technical alert.](https://www.cisa.gov/news-events/cybersecurity-advisories/aa17-181a) 11. [ 7 ][Cybersecurity and Infrastructure Security Agency and Federal Bureau of Investigation. (2023). CL0P ransomware gang exploits MOVEit vulnerability.](https://www.cisa.gov/sites/default/files/2023-06/aa23-158a-stopransomware-cl0p-ransomware-gang-exploits-moveit-vulnerability_2.pdf) 12. [ 8 ][Cybersecurity and Infrastructure Security Agency. Known Exploited Vulnerabilities catalog.](https://www.cisa.gov/known-exploited-vulnerabilities-catalog) Continue exploring ## Complementary ontologies [416 categories NOSIBLE Event Ontology Follow attack behavior into outages, lawsuits, fraud, remediation and other corporate consequences. Explore ontology →](https://nosible.com/ontologies/nosible-events) [237 categories GICS Industry Classification Measure which industries carry the greatest exposure to specific attack tactics and techniques. Explore ontology →](https://nosible.com/ontologies/gics) > Explore 866 MITRE ATT&CK categories and example usages separating ransomware persistence, exfiltration and operational impact across major cyber incidents. **URL:** https://nosible.com/ontologies/mitre-attack --- --- title: "NOSIBLE Event Ontology" description: "Explore 416 corporate event categories and example usages following Credit Suisse, Microsoft-Activision and CrowdStrike through changing risk mechanisms." url: "https://nosible.com/ontologies/nosible-events" --- [Home](https://nosible.com/) [Ontologies](https://nosible.com/ontologies) NOSIBLE Events Ontology field guide By NOSIBLE Research Updated 2026-07-18 # NOSIBLE Event Ontology The NOSIBLE Event Ontology tracks how one corporate shock becomes operational, customer, legal, governance and financial events. Corporate shocks propagate through distinct events. A software failure causes disruption. Remediation triggers customer actions. Losses bring claims and litigation. Management responses create governance, disclosure and earnings events. One topic label cannot preserve that sequence. [[ 1 ]](https://doi.org/10.1016/j.jbankfin.2005.09.015) The NOSIBLE Event Ontology maps reporting into 17 categories, 69 subcategories and 330 event types. World applies those paths to canonical event clusters with point-in-time dates and source coverage. Researchers can construct event chains, separate the initial shock from later consequences and test which stages affect exposed assets across the corporate resolution sequence through time. [[ 2 ]](https://www.finma.ch/en/news/2023/03/20230319-mm-cs-ubs/) [[ 3 ]](https://www.snb.ch/en/publications/communication/press-releases/2023/pre_20230319_1) Categories 416 categories Structure 3 levels World events labelled 29.8% Stable codes Since v1 Release data CC0 On this page [01 Foundations](https://nosible.com/ontologies/nosible-events#foundations) [02 Categories](https://nosible.com/ontologies/nosible-events#vocabulary) [03 Example usages](https://nosible.com/ontologies/nosible-events#trends) [04 Downloads and references](https://nosible.com/ontologies/nosible-events#downloads) Foundations ## The NOSIBLE Event Ontology separates corporate shocks from their downstream consequences The ontology is designed around investable corporate events. Categories separate operational, legal, regulatory, governance, earnings, capital-market, credit, reputation and competitive effects. Subcategories narrow the mechanism. Concrete event types specify the reported action or outcome. The hierarchy supports both broad samples and precise event studies across linked corporate event chains and market regimes. World classifies what a canonical event record reports. It does not prove causality, materiality or legal liability. A cascade can contain repeated reporting, parallel consequences and classification errors. Researchers should preserve event time, deduplicate source clusters, audit leaf labels and test whether downstream events add information beyond the initial shock, its immediate coverage and concurrent market conditions. [[ 1 ]](https://doi.org/10.1016/j.jbankfin.2005.09.015) Categories ## Seventeen categories expand into 330 concrete corporate events across three levels The NOSIBLE Event Ontology contains 416 categories: 17 roots, 69 subcategories and 330 leaf events at maximum depth two. The explorer loads bounded branches and exposes definitions, full paths and World counts. World V1.2 assigns a NOSIBLE path to 4,562,361 of 15,311,040 events, giving 29.8% corpus coverage. Search the ontology 17 / 416 shown World V1.2 coverage Coverage **29.8%** Labelled events **4,562,361** World events **15,311,040** Ontology index 4,562,361 of 15,311,040 World events carry NOSIBLE Event Ontology labels. Select a category to inspect the evidence Hierarchy category / subcategory / event › Business Operations Workforce, production and delivery · 6 children › Capital Markets Financing, issuance and transaction activity · 8 children › Capital Returns Dividends, repurchases and distributions · 4 children › Competitive Dynamics Market structure, rivalry and partnership changes · 1 children › Corporate Governance Leadership, board and control changes · 7 children › Credit Analysis Default, restructuring and creditor outcomes · 3 children › Earnings Insights Reported results, guidance and forecast revisions · 3 children › Fundamental Analysis Analyst estimates, valuation and operating fundamentals · 5 children › Insider Analysis Insider ownership and transaction events · 2 children › Investor Relations Shareholder communication and sentiment · 3 children › Legal Actions Litigation, claims, settlements and court outcomes · 7 children › Public Relations Crisis response and external communication · 2 children › Regulatory Compliance Regulator investigations, approvals and enforcement · 3 children › Reputational Risk Trust, safety, security and stakeholder perception · 6 children › Sustainability Environmental and social operating effects · 2 children › Symbology Changes Ticker, listing and security-identifier changes · 2 children › Technical Analysis Price, volume and market-microstructure events · 5 children Business Operations ### Business Operations Copy link ↗ #### Definition Changes to how an organisation operates, including continuity, workforce, production, sales and delivery. Research use Workforce, production and delivery Code Business Operations Events with label 1,126,976 Share of labelled 24.7% Share of World 7.4% **Events with label** is the selected count. **Label share** divides it by 4,562,361 assigned events; **World share** divides it by all 15,311,040 events. #### Representative World V1.2 events One strong classified example per available year, with up to ten years shown. 10 examples 1. 2015-04-14Microsoft Acquires Datazen to Enhance Mobile Business Intelligence CapabilitiesCoverage 27 2. 2016-10-24Dangote Group Fires 48 Staff Including 36 Expats Amid RecessionCoverage 22 3. 2017-07-18Google Glass Enterprise Edition Returns for Factory and Business UseCoverage 57 4. 2019-07-23Huawei Cuts Over 600 US Jobs Following Blacklisting and SanctionsCoverage 71 5. 2020-03-25Kentucky Governor Halts Evictions and Cuts Non-Essential Business OperationsCoverage 104 6. 2021-03-17Samsung Skips Galaxy Note 21 Due to Global Chip Shortage CrisisCoverage 184 7. 2022-03-02Ford Suspends Russian Operations Citing Ukraine Invasion ConcernsCoverage 223 8. 2024-12-11TikTok Challenges Canadian Order to Shut Down Business OperationsCoverage 193 9. 2025-10-02P&G Exits Pakistan Direct Operations, Shifts to Distributor ModelCoverage 46 10. 2026-05-09XTrend Lab Builds Future of Connected Business OperationsCoverage 139 Browse categories to compare definitions, World V1.2 statistics and representative classified events. Example usage 1 of 3 ### Example Usage: Credit Suisse coverage moves from liquidity crisis toward takeover, governance and resolution The cohort requires Credit Suisse plus crisis, loss, liquidity, deposit, UBS or rescue language. Normalized coverage rises from 91 to 802 events per million after the 19 March 2023 rescue announcement. The fixed entity and risk filter separates a resolution shock from the changing volume of the wider World corpus. [[ 2 ]](https://www.finma.ch/en/news/2023/03/20230319-mm-cs-ubs/) [[ 3 ]](https://www.snb.ch/en/publications/communication/press-releases/2023/pre_20230319_1) [[ 4 ]](https://www.news.admin.ch/en/nsb?id=93793) Capital Markets rises from 14.7% to 23.7% of ontology-assigned events. Corporate Governance rises from 13.2% to 20.0%, while Technical Analysis falls from 17.6% to 4.8%. FINMA and the Swiss National Bank records confirm the transition from confidence and liquidity stress toward takeover, governance and resolution. [[ 2 ]](https://www.finma.ch/en/news/2023/03/20230319-mm-cs-ubs/) [[ 5 ]](https://www.ubs.com/global/en/media/display-page-ndp/en-20230612-ubs-credit-suisse-acquisition.html) Normalized coverage **91 to 802** Capital Markets **14.7% to 23.7%** Corporate Governance **13.2% to 20.0%** **Coverage surges as Credit Suisse reporting moves into capital markets and governance** stacked rate per million World events · 14-day memory Capital Markets Corporate Governance Technical Analysis Credit Analysis Other The cohort requires Credit Suisse plus crisis, loss, liquidity, deposit, UBS or rescue language. The comparison splits at 19 March 2023. Coverage divides matched events by all World events in each phase. Composition divides by ontology-assigned cohort events. The result describes the resolution sequence; it does not measure loss severity, depositor behaviour, counterparty exposure or causal market impact. Example usage 2 of 3 ### Example Usage: Microsoft–Activision moves from antitrust review to layoffs within 100 days of closing Microsoft–Activision reporting changes event type three times. Antitrust Case Filed and Antitrust Case Updated dominate before the 13 October 2023 close. Merger Concluded peaks at completion. Workforce Reduced appears in late January 2024 and becomes the leading event type during the first restructuring wave roughly 100 days after the deal closes. [[ 8 ]](https://www.ftc.gov/legal-library/browse/cases-proceedings/2210077-microsoftactivision-blizzard-matter) Those phases carry different financial questions. Antitrust events change closing probability. Completion starts the ownership and accounting clock. Workforce reductions reveal integration costs and synergy execution. A single acquisition keyword merges all three. Event classifications let researchers align spreads, revisions, operating margins and competitor responses to the mechanism that actually changed during each distinct deal phase. [[ 9 ]](https://blogs.microsoft.com/blog/2023/10/13/welcoming-the-legendary-teams-at-activision-blizzard-king-to-team-xbox/) Corporate Governance +18.9pp Business Operations +9.9pp Legal Actions -14.0pp Regulatory Compliance -16.5pp Twenty-one-day normalized event-type attention · per million Workforce reduced Antitrust filed Merger concluded Antitrust updated The fixed English cohort covers June 2023 through February 2024 and requires Microsoft plus Activision, Blizzard or Call of Duty. Each line pools event-type counts across the current and prior twenty days, then divides by all English World events in that window. Example usage 3 of 3 ### Example Usage: CrowdStrike reporting moves from outage disclosure into fraud cases and apologies CrowdStrike reporting changes event type after the 19 July 2024 outage. Cyber-attack and software-vulnerability classifications peak during the acute response. Fraud Case Filed appears in the following weeks, while Apology Statement Issued arrives later. The line chart exposes the sequence that pooled before-and-after shares conceal. [[ 6 ]](https://www.crowdstrike.com/en-us/blog/falcon-content-update-preliminary-post-incident-report/) Each phase starts a new risk clock. The outage affects clients, vendors and operations at once. Fraud cases and apologies shift attention to legal, insurance, churn and trust costs. Researchers can test prices, sales and client events by phase instead of pooling all CrowdStrike news into one risk signal. [[ 7 ]](https://www.sec.gov/Archives/edgar/data/1535527/000153552725000009/crwd-20250131.htm) Legal Actions +24.4pp Reputational Risk -46.4pp Business Operations +12.2pp Normalized coverage -994 per million Fourteen-day normalized event-type attention · per million Fraud case Apology Cyber attack Software vulnerability The explicit English CrowdStrike-outage cohort covers July through October 2024. Each line pools NOSIBLE event-type counts across the current and prior thirteen days, then divides by all English World events in that window. The chart describes the reporting sequence, not realized losses. Data and sources ## Download NOSIBLE Events and References ### Downloads Download categories and counts as CSV, or the complete machine-readable release as JSON. [**CSV** Definitions and counts Download ↓](https://nosible.com/data/world-v1.2/ontologies/nosible-events/v1/nodes.csv) [**JSON** Full machine-readable release Download ↓](https://nosible.com/data/world-v1.2/ontologies/nosible-events/v1/statistics.json) Release details and citation files [Manifest](https://nosible.com/data/world-v1.2/ontologies/nosible-events/v1/manifest.json) [Release README](https://nosible.com/data/world-v1.2/ontologies/nosible-events/v1/README.md) [CC0 license and scope](https://nosible.com/data/world-v1.2/ontologies/nosible-events/v1/LICENSE.md) [Citation file](https://nosible.com/data/world-v1.2/ontologies/nosible-events/v1/CITATION.cff) [Changelog](https://nosible.com/data/world-v1.2/ontologies/nosible-events/v1/CHANGELOG.md) Need the complete World event schema? [Open the World data dictionary.](https://nosible.com/data-dictionaries#world) ### References 1. [ 1 ][Cummins, J. D., Lewis, C. M., & Wei, R. (2006). The market value impact of operational loss events for US banks and insurers. Journal of Banking & Finance, 30(10), 2605-2634.](https://doi.org/10.1016/j.jbankfin.2005.09.015) 2. [ 2 ][Swiss Financial Market Supervisory Authority. (2023). FINMA approves merger of UBS and Credit Suisse.](https://www.finma.ch/en/news/2023/03/20230319-mm-cs-ubs/) 3. [ 3 ][Swiss National Bank. (2023). Swiss National Bank provides substantial liquidity assistance to support UBS takeover of Credit Suisse.](https://www.snb.ch/en/publications/communication/press-releases/2023/pre_20230319_1) 4. [ 4 ][Swiss Federal Council. (2023). Safeguarding financial market stability: Federal Council welcomes and supports UBS takeover of Credit Suisse.](https://www.news.admin.ch/en/nsb?id=93793) 5. [ 5 ][UBS Group AG. (2023). UBS completes Credit Suisse acquisition.](https://www.ubs.com/global/en/media/display-page-ndp/en-20230612-ubs-credit-suisse-acquisition.html) 6. [ 8 ][U.S. Federal Trade Commission. Microsoft/Activision Blizzard matter, docket 9412.](https://www.ftc.gov/legal-library/browse/cases-proceedings/2210077-microsoftactivision-blizzard-matter) 7. [ 9 ][Microsoft. (2023). Welcoming the teams at Activision Blizzard King to Xbox.](https://blogs.microsoft.com/blog/2023/10/13/welcoming-the-legendary-teams-at-activision-blizzard-king-to-team-xbox/) 8. [ 6 ][CrowdStrike. (2024). Preliminary post-incident review: Falcon content configuration update.](https://www.crowdstrike.com/en-us/blog/falcon-content-update-preliminary-post-incident-report/) 9. [ 7 ][CrowdStrike Holdings, Inc. (2025). Form 10-K for the year ended January 31, 2025.](https://www.sec.gov/Archives/edgar/data/1535527/000153552725000009/crwd-20250131.htm) Continue exploring ## Complementary ontologies [237 categories GICS Industry Classification Aggregate company events into industries to measure where an event mechanism is concentrating. Explore ontology →](https://nosible.com/ontologies/gics) [39 categories Asset Class Ontology Connect corporate event mechanisms to the market channels most likely to absorb them. Explore ontology →](https://nosible.com/ontologies/asset-classes) > Explore 416 corporate event categories and example usages following Credit Suisse, Microsoft-Activision and CrowdStrike through changing risk mechanisms. **URL:** https://nosible.com/ontologies/nosible-events --- --- title: "PLOVER Political Events" description: "Explore 88 PLOVER political-event categories and example usages separating threats, sanctions, strikes and concessions across geopolitical and policy episodes." url: "https://nosible.com/ontologies/plover" --- [Home](https://nosible.com/) [Ontologies](https://nosible.com/ontologies) PLOVER Ontology field guide By NOSIBLE Research Updated 2026-07-18 # PLOVER Political Events PLOVER converts political interactions into comparable event classes, actions and contextual qualifiers through time across markets. Political risk changes through actions, not sentiment alone. Consultation, threats, sanctions, mobilisation and assault represent different interactions with different economic channels. PLOVER standardises those actions into event classes, event-specific modes and independent contexts. [[ 1 ]](https://github.com/openeventdata/PLOVER/blob/master/PLOVER_MANUAL.pdf) [[ 2 ]](https://doi.org/10.2307/2010532) World applies PLOVER to canonical events so researchers can measure escalation, support, coercion and de-escalation through time. Those features can be joined point in time to exposed firms, currencies, commodities and sovereign risk. This page provides the complete structure, World counts, a Russia-Ukraine sequence and reproducible data for testing political actions against market exposures. [[ 3 ]](https://doi.org/10.1257/aer.20191823) [[ 4 ]](https://doi.org/10.20955/wp.2022.032) Categories 88 categories Structure 2 levels World events labelled 11.3% Stable codes Since v1 Release data CC0 On this page [01 Foundations](https://nosible.com/ontologies/plover#foundations) [02 Categories](https://nosible.com/ontologies/plover#vocabulary) [03 Example usages](https://nosible.com/ontologies/plover#trends) [04 Downloads and references](https://nosible.com/ontologies/plover#downloads) Foundations ## PLOVER records political actions while leaving severity and escalation unranked by design PLOVER defines 16 political event classes. Modes refine an event only where the specification permits them. Contexts are independent qualifiers that can accompany any class. The structure preserves the difference between verbal pressure, material coercion, support, consultation and armed action without forcing every episode onto one continuous score. [[ 1 ]](https://github.com/openeventdata/PLOVER/blob/master/PLOVER_MANUAL.pdf) World classifies the action described in an event record. It does not identify a causal actor-target pair, verify intent or reconstruct every interaction in a campaign. A report can contain several actions while the event receives one selected path. Researchers should test class composition and transitions before estimating any scalar escalation measure. [[ 2 ]](https://doi.org/10.2307/2010532) Categories ## Sixteen event classes combine with conditional modes and independent political contexts PLOVER contains 88 categories: 16 event classes, 35 event-specific modes and 37 independent contexts. Event classes and contexts are roots; allowed modes sit beneath their event. The explorer preserves those rules and exposes definitions, paths and World V1.2 counts. World assigns PLOVER to 1,731,529 events, or 11.3% of the corpus. [[ 1 ]](https://github.com/openeventdata/PLOVER/blob/master/PLOVER_MANUAL.pdf) Search the ontology 53 / 88 shown World V1.2 coverage Coverage **11.3%** Labelled events **1,731,529** World events **15,311,040** Ontology index 1,731,529 of 15,311,040 World events carry PLOVER Political Events labels. Select a category to inspect the evidence Hierarchy event / mode / context AGREE Declared consensus or completed agreement › CONSULT Dialogue and policy coordination · 4 children SUPPORT Political or material backing CONCEDE Abandonment of a prior demand COOPERATE Sustained joint action AID Transfer of resources or expertise › RETREAT Withdrawal from a position or engagement · 6 children › REQUEST Explicit call for another actor to act · 4 children › ACCUSE Attribution of fault or wrongdoing · 3 children › REJECT Refusal of a proposal or demand · 4 children › THREATEN Declared prospective harm or coercion · 6 children PROTEST Public opposition and collective dissent › SANCTION Economic or administrative restriction · 4 children › MOBILIZE Preparation or deployment of organised capacity · 4 children COERCE Material compulsion short of assault ASSAULT Direct physical attack asylum AGREE ### AGREE Copy link ↗ #### Definition Political actors reach, ratify or publicly commit to an agreement. Research use Declared consensus or completed agreement Code AGREE Events with label 37,295 Share of labelled 2.2% Share of World 0.2% **Events with label** is the selected count. **Label share** divides it by 1,731,529 assigned events; **World share** divides it by all 15,311,040 events. #### Representative World V1.2 events One strong classified example per available year, with up to ten years shown. 10 examples 1. 2015-11-02China France Agree on Climate Compliance Checks for Paris DealCoverage 22 2. 2016-02-17Saudi Arabia Russia Qatar Venezuela Agree to Freeze Oil OutputCoverage 126 3. 2017-11-19Gujarat Congress and PAAS agree on Patel reservation ahead of pollsCoverage 78 4. 2019-06-10G20 Ministers Agree to Finalize Global Digital Tax Rules by 2020Coverage 27 5. 2020-07-20NFL and Players Union Agree on Daily COVID-19 Testing ProtocolsCoverage 215 6. 2021-10-06Biden and Xi Agree to Abide by Taiwan Agreement Amid Rising TensionsCoverage 84 7. 2022-11-25India and GCC agree to resume free trade agreement negotiationsCoverage 54 8. 2024-10-21India China Agree on LAC Patrolling to Enable Disengagement in LadakhCoverage 748 9. 2025-05-16NATO Members Agree to 5% GDP Defense Spending Target by June SummitCoverage 80 10. 2026-06-13US and Iran Agree to Peace Deal Wording to End WarCoverage 6,631 Browse categories to compare definitions, World V1.2 statistics and representative classified events. Example usage 1 of 3 ### Example Usage: Russia–Ukraine invasion raises sanctions as threats fall during the opening phase The cohort requires a Russian and Ukrainian actor in the same canonical event. Normalized coverage rises from 873 events per million World events before 24 February 2022 to 25,515 during the opening invasion. This controls for the changing size of the World corpus and isolates the relative reporting shock. [[ 6 ]](https://digitallibrary.un.org/record/4051033/files/A_78_PV.56-EN.pdf) The opening invasion changes the action mix. Sanction rises from 9.8% to 18.1% of assigned events, while Threaten falls from 20.7% to 12.1%. Mobilize falls from 19.4% to 10.1%, while Assault rises from 0.4% to 2.2%. The labels distinguish implemented pressure and material action from pre-invasion rhetoric. [[ 4 ]](https://doi.org/10.20955/wp.2022.032) [[ 5 ]](https://doi.org/10.1016/j.gfj.2023.100925) Normalized coverage **873 to 25515** Sanction share **9.8% to 18.1%** Threat share **20.7% to 12.1%** **Sanctions rise as threats fall after Russia invades Ukraine** 60-day memory · share of assigned events sanction threaten mobilize assault The cohort requires a Russian and Ukrainian actor in the same canonical event. Coverage divides those events by all World events in each phase. Action shares divide by PLOVER-assigned cohort events. The opening invasion raises normalized coverage from 873 to 25,515 per million events. Sanction rises from 9.8% to 18.1%, while Threaten falls from 20.7% to 12.1%. The comparison describes the event record; it does not measure battlefield intensity. Example usage 2 of 3 ### Example Usage: UAW action labels clearly separate threatened disruption from actual strike activity United Auto Workers reporting changes when the stand-up strike begins on 15 September 2023. Threaten accounts for 47.6% of assigned actions before the walkout and falls to 9.1% during the strike. Protest rises from 14.3% to 39.4%, separating bargaining pressure from realized action across Ford, General Motors and Stellantis. [[ 7 ]](https://uaw.org/uaw-president-shawn-fain-announces-stand-up-strike-begin-three-locations/) The transition creates a production-risk clock. Threats affect strike probability and contingency planning; Protest marks active disruption to plants, suppliers and inventories. PLOVER aligns automakers and exposed suppliers to reported action before tests of deliveries, margins, revisions or relative returns. [[ 1 ]](https://github.com/openeventdata/PLOVER/blob/master/PLOVER_MANUAL.pdf) Threaten -38.5pp Protest +25.1pp Mobilize +4.3pp Normalized coverage +67 per million Twenty-eight-day pooled share of PLOVER-assigned UAW events · % Threaten Protest Mobilize The fixed English cohort requires UAW or United Auto Workers plus an automaker, strike, contract or wage term from 15 August through 15 December 2023. The chart shows 29 August through 20 November. Lines pool action counts across the current and prior 27 days, then divide by all PLOVER-assigned cohort events across the full contract-negotiation and strike window for direct phase comparison. Example usage 3 of 3 ### Example Usage: Concessions replace threats after the 2023 debt-ceiling agreement reduces immediate default risk The 2023 United States debt-ceiling cohort changes action class when an agreement becomes credible on 27 May. Threaten accounts for 56.2% of 89 negotiation-phase events and falls to 17.6% around agreement. Concede rises from 7.9% to 58.8%, identifying compromise before the Fiscal Responsibility Act is signed. [[ 8 ]](https://www.whitehouse.gov/briefing-room/statements-releases/2023/05/27/statement-from-president-joe-biden-on-bipartisan-budget-agreement/) That distinction separates default-risk bargaining from resolution. A debt-ceiling keyword remains present in both phases while the reported action changes. PLOVER creates a point-in-time clock for testing Treasury-bill dislocations, yields, volatility and cross-asset hedges without using the final agreement to relabel prior observations. [[ 1 ]](https://github.com/openeventdata/PLOVER/blob/master/PLOVER_MANUAL.pdf) Threaten -38.6pp Concede +50.9pp Reject -13.2pp Twenty-one-day pooled share of PLOVER-assigned debt-ceiling events · % Concede Threaten Reject The fixed English cohort requires United States debt-ceiling or debt-limit language from April through July 2023. The chart shows 1 May through 15 June. Lines pool PLOVER action counts across the current and prior 20 days, then divide by all assigned cohort events. The agreement window contains 17 assigned events, so the transition remains descriptive rather than a causal market result from broad news volume. Data and sources ## Download PLOVER and References ### Downloads Download categories and counts as CSV, or the complete machine-readable release as JSON. [**CSV** Definitions and counts Download ↓](https://nosible.com/data/world-v1.2/ontologies/plover/v1/nodes.csv) [**JSON** Full machine-readable release Download ↓](https://nosible.com/data/world-v1.2/ontologies/plover/v1/statistics.json) Release details and citation files [Manifest](https://nosible.com/data/world-v1.2/ontologies/plover/v1/manifest.json) [Release README](https://nosible.com/data/world-v1.2/ontologies/plover/v1/README.md) [CC0 license and scope](https://nosible.com/data/world-v1.2/ontologies/plover/v1/LICENSE.md) [Citation file](https://nosible.com/data/world-v1.2/ontologies/plover/v1/CITATION.cff) [Changelog](https://nosible.com/data/world-v1.2/ontologies/plover/v1/CHANGELOG.md) Need the complete World event schema? [Open the World data dictionary.](https://nosible.com/data-dictionaries#world) ### References 1. [ 1 ][Open Event Data Alliance. PLOVER: Political Language Ontology for Verifiable Event Records.](https://github.com/openeventdata/PLOVER/blob/master/PLOVER_MANUAL.pdf) 2. [ 2 ][Goldstein, J. S. (1992). A conflict-cooperation scale for WEIS events data. Journal of Conflict Resolution, 36(2), 369-385.](https://doi.org/10.2307/2010532) 3. [ 3 ][Caldara, D., & Iacoviello, M. (2022). Measuring geopolitical risk. American Economic Review, 112(4), 1194-1225.](https://doi.org/10.1257/aer.20191823) 4. [ 4 ][Neely, C. J. (2022). Financial market reactions to the Russian invasion of Ukraine. Federal Reserve Bank of St. Louis Working Paper 2022-032.](https://doi.org/10.20955/wp.2022.032) 5. [ 5 ][Zaremba, A. et al. (2024). Empirical effects of sanctions and support measures on stock prices and exchange rates in the Russia-Ukraine war. Global Finance Journal, 59, 100925.](https://doi.org/10.1016/j.gfj.2023.100925) 6. [ 6 ][United Nations General Assembly. (2024). Official Records, 78th Session, 56th Plenary Meeting. A/78/PV.56.](https://digitallibrary.un.org/record/4051033/files/A_78_PV.56-EN.pdf) 7. [ 7 ][United Auto Workers. (2023). Stand Up Strike begins at three locations.](https://uaw.org/uaw-president-shawn-fain-announces-stand-up-strike-begin-three-locations/) 8. [ 8 ][The White House. (2023). Statement on the bipartisan budget agreement, 27 May 2023.](https://www.whitehouse.gov/briefing-room/statements-releases/2023/05/27/statement-from-president-joe-biden-on-bipartisan-budget-agreement/) Continue exploring ## Complementary ontologies [231 categories World Geography Locate political actions and trace how escalation moves across countries and regions. Explore ontology →](https://nosible.com/ontologies/geography) [186 categories UN Sustainable Development Goals Connect policy action to the economic, social and environmental targets it is intended to change. Explore ontology →](https://nosible.com/ontologies/sustainable-development-goals) > Explore 88 PLOVER political-event categories and example usages separating threats, sanctions, strikes and concessions across geopolitical and policy episodes. **URL:** https://nosible.com/ontologies/plover --- --- title: "Schema.org Event Types" description: "Explore 26 Schema.org event categories and example usages that expose repeatable clocks around company conferences, retail events and entertainment releases." url: "https://nosible.com/ontologies/schema-org-events" --- [Home](https://nosible.com/) [Ontologies](https://nosible.com/ontologies) Schema.org Events Ontology field guide By NOSIBLE Research Updated 2026-07-18 # Schema.org Event Types Schema.org Events identifies event format, allowing scheduled conferences and unscheduled developments to be analysed separately. The date of an information event matters. A planned conference creates anticipation, a fixed disclosure window and a measurable decay path. A delivery, publication, sale or sports event follows a different clock. Schema.org Events supplies compact categories for separating those formats before an event study is specified. [[ 1 ]](https://schema.org/Event) World assigns one of 26 event types to relevant reporting. Researchers can use the field to define scheduled cohorts, align observations to known dates and test whether attention, revisions or price discovery concentrate around the event. The field describes format. It does not identify announcement stage, materiality, surprise or market impact without additional event data, market data and observed returns. Categories 26 categories Structure Flat, no hierarchy World events labelled 6.6% Stable codes Since v1 Release data CC0 On this page [01 Foundations](https://nosible.com/ontologies/schema-org-events#foundations) [02 Categories](https://nosible.com/ontologies/schema-org-events#vocabulary) [03 Example usages](https://nosible.com/ontologies/schema-org-events#trends) [04 Downloads and references](https://nosible.com/ontologies/schema-org-events#downloads) Foundations ## Schema.org Events classifies event format without claiming timing, surprise or materiality Schema.org defines Event as an occurrence at a time and location, with specific types for conferences, business events, publications, broadcasts and other formats. World uses 26 flat, single-label event categories. ConferenceEvent and BusinessEvent are peers. Neither is a parent of the other in the World field. [[ 1 ]](https://schema.org/Event) [[ 2 ]](https://schema.org/ConferenceEvent) [[ 3 ]](https://schema.org/BusinessEvent) The label answers a narrow question: what kind of event is described? It does not encode announcement lifecycle. A GTC keynote can produce conference reporting, product announcements and partner releases. Those records may receive different labels or no Schema.org Event label. Event-time analysis therefore needs an external calendar and an entity-resolved cohort in addition to the ontology and its event-format classification. Categories ## Schema.org Events uses 26 peer categories for scheduled and published event formats The ontology is flat. Every assigned event receives one peer-level label. The explorer shows each category’s definition, stable code and World V1.2 count. Search keeps all formats accessible on one page. EventSeries identifies a recurring group, but the field does not link individual records into a lifecycle or series instance. [[ 1 ]](https://schema.org/Event) Search the ontology 26 / 26 shown World V1.2 coverage Coverage **6.6%** Labelled events **1,010,214** World events **15,311,040** Ontology index 1,010,214 of 15,311,040 World events carry Schema.org Event Types labels. Select a category to inspect the evidence Categories schema_org_event BroadcastEvent Live or scheduled public-media transmission BusinessEvent Commercial briefings, investor sessions and corporate gatherings ChildrensEvent ComedyEvent ConferenceEvent Conferences, summits and multi-session professional gatherings CourseInstance DanceEvent DeliveryEvent Planned arrival or handover of goods EducationEvent Event EventSeries Recurring events sharing one identity or programme ExhibitionEvent Time-bounded product, technology or industry exhibitions Festival FoodEvent Hackathon LiteraryEvent MusicEvent BroadcastEvent ### BroadcastEvent Copy link ↗ #### Definition Reporting about a scheduled live or recorded programme transmitted to a public audience during a defined broadcast window. Research use Live or scheduled public-media transmission Code BroadcastEvent Events with label 20,561 Share of labelled 2.0% Share of World 0.1% **Events with label** is the selected count. **Label share** divides it by 1,010,214 assigned events; **World share** divides it by all 15,311,040 events. #### Representative World V1.2 events One strong classified example per available year, with up to ten years shown. 10 examples 1. 2015-12-03Facebook Launches Public Live Video Streaming For All UsersCoverage 82 2. 2016-07-05PBS Apologizes for Airing Old Fireworks Footage During Live July 4th BroadcastCoverage 33 3. 2017-02-15Two Dominican Journalists Killed During Live Facebook BroadcastCoverage 40 4. 2019-01-27Rent Live Fox Broadcast Proceeds After Actor Brennin Hunt Breaks FootCoverage 57 5. 2020-12-22Biden Receives COVID-19 Vaccine Live on TV to Boost Public ConfidenceCoverage 255 6. 2021-10-28Vatican Cancels Live Broadcast of Biden Meeting Pope FrancisCoverage 273 7. 2022-10-09Iran Protesters Hack State TV During Live Broadcast Amid Ongoing UprisingsCoverage 838 8. 2024-01-09Armed Gang Storms Ecuador TV Station During Live Broadcast Amid Escalating ViolenceCoverage 1,156 9. 2025-06-10CNN Journalist Detained Live on Air During Los Angeles ProtestsCoverage 137 10. 2026-03-06BBC Regrets N-Word Slur Airing During Live BAFTA Awards BroadcastCoverage 106 Browse categories to compare definitions, World V1.2 statistics and representative classified events. Example usage 1 of 3 ### Example Usage: Corporate conferences create repeatable information clocks around each scheduled opening keynote day Ten NVIDIA GTC, Google I/O and AWS re:Invent cycles from 2023 through 2026 produce the same event-time pattern. Nine peak within one day of the opening keynote. Normalized attention rises from 128 events per million English World events during the baseline to 3,669 on keynote day, creating a repeatable disclosure clock. [[ 4 ]](https://nvidianews.nvidia.com/news/see-the-future-at-gtc-2024-nvidias-jensen-huang-to-unveil-latest-breakthroughs-in-accelerated-computing-generative-aiand-robotics) [[ 5 ]](https://nvidianews.nvidia.com/news/nvidia-ceo-jensen-huang-and-industry-visionaries-to-unveil-whats-next-in-ai-at-gtc-2025/) [[ 6 ]](https://nvidianews.nvidia.com/news/nvidia-ceo-jensen-huang-and-global-technology-leaders-to-showcase-age-of-ai-at-gtc-2026) [[ 7 ]](https://blog.google/innovation-and-ai/technology/developers-tools/io-2023/) [[ 8 ]](https://blog.google/innovation-and-ai/technology/developers-tools/google-io-2024-collection/) [[ 9 ]](https://blog.google/innovation-and-ai/technology/developers-tools/google-io-2025-collection/) [[ 10 ]](https://blog.google/innovation-and-ai/technology/developers-tools/google-io-2026-collection/) [[ 11 ]](https://aws.amazon.com/blogs/aws/category/events/reinvent/) Only 13.4% of assigned records use scheduled ConferenceEvent or BusinessEvent formats. That distinction stops surrounding announcements from masquerading as the conference itself. Researchers can align suppliers, customers and competitors to keynote day, then test volatility, revisions, relative returns and adoption without pooling weeks of unrelated company coverage or inferring an announcement lifecycle from labels alone. [[ 12 ]](https://doi.org/10.1111/j.1475-679X.2011.00426.x) NVIDIA GTC / Google I/O / AWS re:Invent / 2023-2026 558 events / 10 conference cycles **Corporate conferences create repeatable information clocks around keynote day** seven-day normalized attention / event day NVIDIA GTC Google I/O AWS re:Invent Each line pools repeated conference cycles by event day and divides classified events by all English World events in the same seven-day window. 9 of 10 cycles peak within one day of the keynote. The recurring clock isolates predictable product and platform disclosure windows from unscheduled company news. Example usage 2 of 3 ### Example Usage: Prime Day creates a repeatable opening-day signal across three expanding retail cycles Prime Day creates a repeatable event-time signature across three annual retail cycles. The normalized Schema.org Event rate peaks on the official opening day in 2023, 2024 and 2025. Amazon expanded the 2025 sale from two days to four, yet the classified attention spike remains anchored to the first day of the event. [[ 13 ]](https://www.aboutamazon.com/news/retail/amazon-prime-day-2024-date) SaleEvent accounts for 96.1% of assigned events in 2023 and 92.2% in 2025. Normalized assigned coverage rises from 1,011 to 1,614 events per million across those cycles. The category isolates a scheduled commercial demand window that ordinary Amazon mentions mix with products, logistics, advertising and corporate reporting. [[ 14 ]](https://www.aboutamazon.com/news/retail/what-makes-prime-day-2025-different) Normalized coverage +603 per million SaleEvent share -3.9pp Assigned events +78 Normalized assigned attention around opening day · per million 2023 2024 2025 The English Prime Day cohort uses matched 29-day windows around official opening dates in 2023, 2024 and 2025. Daily Schema.org Event counts are divided by all English World events; the horizontal axis is event day, with zero marking the opening. Example usage 3 of 3 ### Example Usage: Eras Tour coverage moves from live concerts into cinemas and streaming Eras Tour reporting changes event format as the product moves through distribution channels. MusicEvent accounts for 66.1% of assigned events during the live-tour phase. ScreeningEvent reaches 79.4% around the theatrical release. The same property becomes a new research cohort when distribution moves from venues into cinemas. [[ 15 ]](https://www.amctheatres.com/movies/taylor-swift-the-eras-tour-74500) Streaming creates a third transition. OnDemandEvent reaches 70.6% of assigned coverage in March 2024, when Taylor's Version arrives on Disney+. Schema.org Event categories identify the relevant commercial clock without changing the entity query. Researchers can separate ticket sales, theatrical distribution and subscription engagement without pooling every Eras Tour mention across markets and distribution channels. [[ 16 ]](https://press.disneyplus.com/press.disneyplus.com/news/taylor-swift-the-eras-tour-taylors-version-now-streaming) Normalized coverage +119 per million MusicEvent -63.2pp ScreeningEvent +76.0pp Monthly share of assigned event formats · % MusicEvent ScreeningEvent OnDemandEvent The explicit English Eras Tour cohort spans March 2023 through April 2024. Monthly shares use Schema.org Event-assigned records; normalized phase rates divide assigned events by all English World events in the same months for direct comparison. Data and sources ## Download Schema.org Events and References ### Downloads Download categories and counts as CSV, or the complete machine-readable release as JSON. [**CSV** Definitions and counts Download ↓](https://nosible.com/data/world-v1.2/ontologies/schema-org-events/v1/nodes.csv) [**JSON** Full machine-readable release Download ↓](https://nosible.com/data/world-v1.2/ontologies/schema-org-events/v1/statistics.json) Release details and citation files [Manifest](https://nosible.com/data/world-v1.2/ontologies/schema-org-events/v1/manifest.json) [Release README](https://nosible.com/data/world-v1.2/ontologies/schema-org-events/v1/README.md) [CC0 license and scope](https://nosible.com/data/world-v1.2/ontologies/schema-org-events/v1/LICENSE.md) [Citation file](https://nosible.com/data/world-v1.2/ontologies/schema-org-events/v1/CITATION.cff) [Changelog](https://nosible.com/data/world-v1.2/ontologies/schema-org-events/v1/CHANGELOG.md) Need the complete World event schema? [Open the World data dictionary.](https://nosible.com/data-dictionaries#world) ### References 1. [ 1 ][Schema.org. Event type definition and subtype hierarchy.](https://schema.org/Event) 2. [ 2 ][Schema.org. ConferenceEvent type definition.](https://schema.org/ConferenceEvent) 3. [ 3 ][Schema.org. BusinessEvent type definition.](https://schema.org/BusinessEvent) 4. [ 4 ][NVIDIA. (2024, February 20). See the future at GTC 2024.](https://nvidianews.nvidia.com/news/see-the-future-at-gtc-2024-nvidias-jensen-huang-to-unveil-latest-breakthroughs-in-accelerated-computing-generative-aiand-robotics) 5. [ 5 ][NVIDIA. (2025, March 5). NVIDIA CEO Jensen Huang and industry visionaries to unveil what is next in AI at GTC 2025.](https://nvidianews.nvidia.com/news/nvidia-ceo-jensen-huang-and-industry-visionaries-to-unveil-whats-next-in-ai-at-gtc-2025/) 6. [ 6 ][NVIDIA. (2026, March 3). NVIDIA CEO Jensen Huang and global technology leaders to showcase age of AI at GTC 2026.](https://nvidianews.nvidia.com/news/nvidia-ceo-jensen-huang-and-global-technology-leaders-to-showcase-age-of-ai-at-gtc-2026) 7. [ 7 ][Google. (2023, May 10). Google I/O 2023 announcements.](https://blog.google/innovation-and-ai/technology/developers-tools/io-2023/) 8. [ 8 ][Google. (2024, May 14). Google I/O 2024 announcements.](https://blog.google/innovation-and-ai/technology/developers-tools/google-io-2024-collection/) 9. [ 9 ][Google. (2025, May 20). Google I/O 2025 announcements.](https://blog.google/innovation-and-ai/technology/developers-tools/google-io-2025-collection/) 10. [ 10 ][Google. (2026, May 19). Google I/O 2026 announcements.](https://blog.google/innovation-and-ai/technology/developers-tools/google-io-2026-collection/) 11. [ 11 ][Amazon Web Services. AWS re:Invent official schedules and announcement recaps, 2023–2025.](https://aws.amazon.com/blogs/aws/category/events/reinvent/) 12. [ 12 ][Bushee, B. J., Jung, M. J., & Miller, G. S. (2011). Conference presentations and the disclosure milieu. Journal of Accounting Research, 49(5), 1163–1192.](https://doi.org/10.1111/j.1475-679X.2011.00426.x) 13. [ 13 ][Amazon. (2024). Prime Day returns July 16-17 for its tenth edition.](https://www.aboutamazon.com/news/retail/amazon-prime-day-2024-date) 14. [ 14 ][Amazon. (2025). Prime Day 2025 expands to four days, July 8-11.](https://www.aboutamazon.com/news/retail/what-makes-prime-day-2025-different) 15. [ 15 ][AMC Theatres. Taylor Swift: The Eras Tour theatrical release, 13 October 2023.](https://www.amctheatres.com/movies/taylor-swift-the-eras-tour-74500) 16. [ 16 ][Disney+. (2024). Taylor Swift: The Eras Tour (Taylor's Version) begins streaming 15 March 2024.](https://press.disneyplus.com/press.disneyplus.com/news/taylor-swift-the-eras-tour-taylors-version-now-streaming) Continue exploring ## Complementary ontologies [55 categories IPTC News Genre Add the reporting format to distinguish previews, supplied material, analysis and confirmed results. Explore ontology →](https://nosible.com/ontologies/iptc-genre) [416 categories NOSIBLE Event Ontology Connect generic event shape to financially specific corporate actions and consequences. Explore ontology →](https://nosible.com/ontologies/nosible-events) > Explore 26 Schema.org event categories and example usages that expose repeatable clocks around company conferences, retail events and entertainment releases. **URL:** https://nosible.com/ontologies/schema-org-events --- --- title: "SportsML Categories" description: "Explore 336 SportsML categories and example usages measuring padel growth, Olympic breakdancing and new Formula One attention windows in World V1.2." url: "https://nosible.com/ontologies/sportsml" --- [Home](https://nosible.com/) [Ontologies](https://nosible.com/ontologies) SportsML Ontology field guide By NOSIBLE Research Updated 2026-07-18 # SportsML Categories SportsML separates sport disciplines from competition types so tournament coverage can be compared consistently across seasons. Sports reporting combines different disciplines, competition formats and tournament stages. A World Cup, league match and Olympic event are not interchangeable observations. SportsML supplies controlled categories for those distinctions and supports consistent exchange across publishers and data systems. [[ 1 ]](https://iptc.org/standards/sportsml-g2/) World applies independent sport and event-type labels to canonical events. Researchers can test whether coverage composition improves audience, rights or sponsor models beyond raw article volume. This page provides every category, World counts, an annual padel comparison and reproducible observations. [[ 2 ]](https://www.padelfip.com/2024/05/world-padel-report-2024-the-first-official-fip-report-on-the-padel-movement/) Categories 336 categories Structure Flat, no hierarchy World events labelled 8.3% Stable codes Since v1 Release data CC0 On this page [01 Foundations](https://nosible.com/ontologies/sportsml#foundations) [02 Categories](https://nosible.com/ontologies/sportsml#vocabulary) [03 Example usages](https://nosible.com/ontologies/sportsml#trends) [04 Downloads and references](https://nosible.com/ontologies/sportsml#downloads) Foundations ## SportsML standardises sports and competition types without measuring commercial performance IPTC designed SportsML as an open interchange standard for sports schedules, results, standings and event reports. Its controlled categories keep disciplines and competition concepts stable across providers. World uses an IPTC-aligned subset for sport and event-type classification across the complete event corpus. [[ 1 ]](https://iptc.org/standards/sportsml-g2/) World classifies what the event record reports. It does not identify a complete fixture, audience, rights contract, sponsorship exposure or revenue contribution. Sport and event type are independent fields, not a parent-child path. Researchers should validate both assignment slots and join attention to audited commercial data separately. [[ 2 ]](https://www.padelfip.com/2024/05/world-padel-report-2024-the-first-official-fip-report-on-the-padel-movement/) Categories ## SportsML separates 316 sports categories from 20 independent competition-event types SportsML contains 336 peer-level categories: 316 sport disciplines and 20 competition-event types. World assigns at least one SportsML field to 1,266,796 canonical events, or 8.3% of World V1.2. The explorer exposes each definition and count while keeping the two fields structurally separate for analysis. [[ 1 ]](https://iptc.org/standards/sportsml-g2/) Search the ontology 100 / 336 shown World V1.2 coverage Coverage **8.3%** Labelled events **1,266,796** World events **15,311,040** Ontology index 1,266,796 of 15,311,040 World events carry SportsML Categories labels. Select a category to inspect the evidence Categories sport / event_type 3x3 basketball Three-player basketball events American football Australian rules football C1 (canoeing) (retired) C2 (canoeing) (retired) C4 (canoeing) (retired) Canadian football F3000 (retired) Formula One (retired) Gaelic football Indy Racing (retired) Jai Alai (Pelota) K1 (kayaking) (retired) K2 (kayaking) (retired) K4 (kayaking) (retired) Keirin (retired) Madison race (retired) Load next categories · 236 remaining 3x3 basketball ### 3x3 basketball Copy link ↗ #### Definition Text about 3x3 basketball, the competitive team sport with three players per side on a half-court under a distinct FIBA ruleset for limited-duration games and tournament events. Research use Three-player basketball events Code 3x3 basketball Events with label 10,570 Share of labelled 0.8% Share of World 0.1% **Events with label** is the selected count. **Label share** divides it by 1,266,796 assigned events; **World share** divides it by all 15,311,040 events. #### Representative World V1.2 events One strong classified example per available year, with up to ten years shown. 10 examples 1. 2015-07-13Team USA Wins Men's Basketball Gold at World University GamesCoverage 18 2. 2016-08-30Three Tennessee High School Basketball Players Convicted of RapeCoverage 16 3. 2017-06-09IOC Adds 3x3 Basketball to 2020 Tokyo Olympics for Urban AppealCoverage 78 4. 2019-12-03Three Georgetown Basketball Players Face Restraining Orders for AssaultCoverage 38 5. 2020-03-10Nebraska Adds Two Football Players to Basketball for Big Ten TournamentCoverage 25 6. 2021-07-28US Women Win Historic 3x3 Basketball Gold in Tokyo Olympic DebutCoverage 378 7. 2022-12-21Northwest Boys Basketball Overcome 0-3 Start to Win Three Straight GamesCoverage 36 8. 2024-07-30US 3x3 Women's Basketball Team Loses Opener to Germany in ParisCoverage 130 9. 2025-10-13Jumpshot Singapore Hits World No. 35 in FIBA 3x3 RankingsCoverage 51 10. 2026-04-04Parker, Delle Donne, 1996 Olympic Team Enshrined in Basketball Hall of FameCoverage 538 Browse categories to compare definitions, World V1.2 statistics and representative classified events. Example usage 1 of 3 ### Example Usage: Padel coverage nearly triples as global participation and commercial attention expand From 2019 to 2025, padel rises from 34.3 to 94.9 assignments per 100,000 World events, a 2.77-fold increase after controlling for corpus growth. Its share of SportsML assignments rises from 0.38% to 1.17%, a 3.11-fold increase. Both measures show that padel gains attention within and beyond sport coverage. [[ 2 ]](https://www.padelfip.com/2024/05/world-padel-report-2024-the-first-official-fip-report-on-the-padel-movement/) [[ 3 ]](https://www.padelfip.com/world-padel-report-2025/) Tennis accounts for 4.05% of sport assignments in 2019 and 4.49% in 2025. Those endpoints provide scale; they do not establish a stable path or control group. The commercial question is whether audiences, sponsorship inventory and media-rights prices keep pace with reporting. World supplies the labels, not the outcome. [[ 1 ]](https://iptc.org/standards/sportsml-g2/) [[ 4 ]](https://playtomic.com/global-padel-report) [[ 5 ]](https://premierpadel.com/en/news/premier-padel-announces-groundbreaking-strategic-partnership-with-red-bull) Normalized coverage **34.3 to 94.9** Coverage increase **2.77x** Sports-share increase **3.11x** **Normalized padel coverage rises from 34 to 95 per 100,000 events** padel assignments per 100,000 World events / trailing 12 months The composition chart divides padel and tennis assignments by all SportsML sport assignments in each complete calendar year. The trend chart divides padel assignments by all World events over a trailing 12-month window. Raw counts are denominator diagnostics only. Neither measure proves participation, audiences, rights fees, sponsorship demand, commercial value or investment returns in this World sample. Example usage 2 of 3 ### Example Usage: Olympic breakdancing creates a new classified attention cycle at Paris 2024 Breaking is Olympic breakdancing. It debuted at Paris 2024 and creates a distinct classified attention cycle inside the wider Games. Normalized coverage rises from 339 to 1,272 events per million compared with the matched Tokyo window, a 3.75-fold increase. The daily lines show when the new discipline enters and leaves the news cycle. [[ 6 ]](https://library.olympics.com/cnospa/digitalCollection/DigitalCollectionAttachmentDownloadHandler.ashx?documentId=3415063&parentDocumentId=3415062&skipCopyright=true&skipWatermark=true) That separation matters to sponsors, broadcasters and consumer brands evaluating a new sport rather than the Olympics as one aggregate property. SportsML creates comparable discipline-level clocks for attention, endorsement and audience tests. The method separates a genuine launch from host-city, medal-table and corpus-growth effects. [[ 6 ]](https://library.olympics.com/cnospa/digitalCollection/DigitalCollectionAttachmentDownloadHandler.ashx?documentId=3415063&parentDocumentId=3415062&skipCopyright=true&skipWatermark=true) Breaking +933 per million 3x3 basketball +201 per million Sport climbing -2 per million Skateboarding -808 per million Daily normalized breakdancing attention · per million Paris 2024 Tokyo 2020 The matched Tokyo and Paris Olympic windows contain 266 and 317 SportsML-assigned events. Each sport count is divided by all English World events in its window, so the comparison controls for corpus growth rather than relying on raw event counts. Example usage 3 of 3 ### Example Usage: Miami and Las Vegas add two new Formula One attention windows The United States Formula One calendar expands from Austin alone in 2019 to Miami, Austin and Las Vegas in 2024. Fourteen-day normalized SportsML attention produces one late-year peak in 2019 and three distinct peaks in 2024. Classified race-window coverage rises from 30 to 126 events per million across the wider comparison. [[ 7 ]](https://www.formula1.com/en/racing/2022/miami) Each peak defines a commercial clock for airlines, casinos, hotels, sponsors, broadcasters and host cities. SportsML separates Formula One from other racing and sports news. Researchers can compare local exposure, sponsorship and attention decay across races instead of pooling an expanded calendar into one yearly observation. [[ 8 ]](https://www.formula1.com/en/racing/2023/las-vegas) Normalized coverage +95.2 per million Scheduled US races +2 Fourteen-day normalized US Formula One attention by calendar date · per million 2019 · Austin only 2024 · three races The fixed English cohorts cover April through mid-December 2019 and 2024. Formula One and motor-car-racing assignments naming a United States, Austin, Miami or Las Vegas Grand Prix are pooled across the current and prior 13 days, then divided by all English World events in the same window. Data and sources ## Download SportsML and References ### Downloads Download categories and counts as CSV, or the complete machine-readable release as JSON. [**CSV** Definitions and counts Download ↓](https://nosible.com/data/world-v1.2/ontologies/sportsml/v1/nodes.csv) [**JSON** Full machine-readable release Download ↓](https://nosible.com/data/world-v1.2/ontologies/sportsml/v1/statistics.json) Release details and citation files [Manifest](https://nosible.com/data/world-v1.2/ontologies/sportsml/v1/manifest.json) [Release README](https://nosible.com/data/world-v1.2/ontologies/sportsml/v1/README.md) [CC0 license and scope](https://nosible.com/data/world-v1.2/ontologies/sportsml/v1/LICENSE.md) [Citation file](https://nosible.com/data/world-v1.2/ontologies/sportsml/v1/CITATION.cff) [Changelog](https://nosible.com/data/world-v1.2/ontologies/sportsml/v1/CHANGELOG.md) Need the complete World event schema? [Open the World data dictionary.](https://nosible.com/data-dictionaries#world) ### References 1. [ 1 ][International Press Telecommunications Council. SportsML: The open standard for sports data. SportsML-G2 3.1.](https://iptc.org/standards/sportsml-g2/) 2. [ 2 ][International Padel Federation. World Padel Report 2024.](https://www.padelfip.com/2024/05/world-padel-report-2024-the-first-official-fip-report-on-the-padel-movement/) 3. [ 3 ][International Padel Federation. World Padel Report 2025.](https://www.padelfip.com/world-padel-report-2025/) 4. [ 4 ][Playtomic and Strategy&. Global Padel Report 2025.](https://playtomic.com/global-padel-report) 5. [ 5 ][Premier Padel. Multi-year media, streaming and sponsorship partnership with Red Bull.](https://premierpadel.com/en/news/premier-padel-announces-groundbreaking-strategic-partnership-with-red-bull) 6. [ 6 ][Paris 2024 Organising Committee. Official report of the Olympic and Paralympic Games Paris 2024.](https://library.olympics.com/cnospa/digitalCollection/DigitalCollectionAttachmentDownloadHandler.ashx?documentId=3415063&parentDocumentId=3415062&skipCopyright=true&skipWatermark=true) 7. [ 7 ][Formula One. (2022). Miami Grand Prix official race page.](https://www.formula1.com/en/racing/2022/miami) 8. [ 8 ][Formula One. (2023). Las Vegas Grand Prix official race page.](https://www.formula1.com/en/racing/2023/las-vegas) Continue exploring ## Complementary ontologies [231 categories World Geography Locate where sports attention emerges and compare host markets on a consistent geographic hierarchy. Explore ontology →](https://nosible.com/ontologies/geography) [26 categories Schema.org Event Types Add event format and schedule to competition categories for cleaner recurring-event clocks. Explore ontology →](https://nosible.com/ontologies/schema-org-events) > Explore 336 SportsML categories and example usages measuring padel growth, Olympic breakdancing and new Formula One attention windows in World V1.2. **URL:** https://nosible.com/ontologies/sportsml --- --- title: "UN Sustainable Development Goals" description: "Explore 186 UN Sustainable Development Goal categories and example usages on climate finance, clean-energy implementation and European gas-demand reduction." url: "https://nosible.com/ontologies/sustainable-development-goals" --- [Home](https://nosible.com/) [Ontologies](https://nosible.com/ontologies) UN SDGs Ontology field guide By NOSIBLE Research Updated 2026-07-18 # UN Sustainable Development Goals The UN Sustainable Development Goals separate climate reporting into policy, energy, infrastructure, cities and finance cohorts. Climate policy does not affect one market. It changes power systems, transport, buildings, industry, public finance and development funding. A single climate label collapses those channels. The Sustainable Development Goals preserve them through 17 goals and 169 targets, allowing researchers to separate broad policy attention from the specific implementation mechanism under discussion across markets. [[ 1 ]](https://digitallibrary.un.org/record/3923923?ln=en) [[ 2 ]](https://sdgs.un.org/goals/goal13) World assigns one goal and one target to each classified event. Goal-level data support cross-sector attention measures. Targets isolate investable mechanisms such as renewable energy, resilient infrastructure, urban planning, climate policy and international finance. The labels organise reporting. They do not measure policy delivery, capital committed, emissions avoided or subsequent market repricing. [[ 7 ]](https://www.ipcc.ch/report/ar6/wg3/chapter/chapter-15/) Categories 186 categories Structure 2 levels World events labelled 28.1% Stable codes Since v2 Release data CC0 On this page [01 Foundations](https://nosible.com/ontologies/sustainable-development-goals#foundations) [02 Categories](https://nosible.com/ontologies/sustainable-development-goals#vocabulary) [03 Example usages](https://nosible.com/ontologies/sustainable-development-goals#trends) [04 Downloads and references](https://nosible.com/ontologies/sustainable-development-goals#downloads) Foundations ## The goals connect social and environmental outcomes to measurable implementation mechanisms The 2030 Agenda defines 17 integrated goals. Each goal contains outcome targets and explicit means-of-implementation targets. The distinction matters. Climate Action can describe resilience or policy integration, while target 13.a isolates international climate-finance commitments. Affordable and Clean Energy separates access, renewables, efficiency, cooperation and infrastructure across different investment channels. [[ 1 ]](https://digitallibrary.un.org/record/3923923?ln=en) [[ 2 ]](https://sdgs.un.org/goals/goal13) The goals are policy categories, not mutually exclusive economic sectors. World stores the dominant goal and target for each event. Secondary themes are not retained in this field. Researchers should therefore inspect neighbouring goals and target-level assignments before interpreting a change as a shift in the real economy. [[ 7 ]](https://www.ipcc.ch/report/ar6/wg3/chapter/chapter-15/) Categories ## Seventeen goals and 169 targets separate policy outcomes from investment mechanisms Seventeen roots define the goals. Their 169 children define outcomes and implementation mechanisms. Letter-suffixed targets identify means of implementation; numeric targets identify outcomes. The explorer exposes every definition, code, path and World V1.2 count without creating a separate page for each target. [[ 1 ]](https://digitallibrary.un.org/record/3923923?ln=en) Search the ontology 17 / 186 shown World V1.2 coverage Coverage **28.1%** Labelled events **4,301,492** World events **15,311,040** Ontology index 4,301,492 of 15,311,040 World events carry UN Sustainable Development Goals labels. Select a category to inspect the evidence Hierarchy goal / target › No Poverty Poverty reduction and social protection · 7 children › Zero Hunger · 8 children › Good Health and Well-Being · 13 children › Quality Education · 10 children › Gender Equality · 9 children › Clean Water and Sanitation · 8 children › Affordable and Clean Energy Energy access, renewables, efficiency and clean-energy infrastructure · 5 children › Decent Work and Economic Growth · 12 children › Industry, Innovation and Infrastructure Resilient infrastructure, industrial transition and technology · 8 children › Reduced Inequalities · 10 children › Sustainable Cities and Communities Buildings, transport, urban resilience and disaster exposure · 10 children › Responsible Consumption and Production · 11 children › Climate Action Physical resilience, climate policy, capability and finance · 5 children › Life Below Water · 10 children › Life on Land · 12 children › Peace, Justice and Strong Institutions · 12 children › Partnerships for the Goals Finance, debt, technology transfer, trade and institutional cooperation · 19 children No Poverty ### No Poverty Copy link ↗ #### Definition Reporting on poverty reduction, social protection and resilience. Research use Poverty reduction and social protection Code No Poverty Events with label 34,955 Share of labelled 0.8% Share of World 0.2% **Events with label** is the selected count. **Label share** divides it by 4,301,492 assigned events; **World share** divides it by all 15,311,040 events. #### Representative World V1.2 events One strong classified example per available year, with up to ten years shown. 10 examples 1. 2015-10-12Angus Deaton Wins 2015 Nobel Prize for Economics on PovertyCoverage 337 2. 2016-10-03World Bank warns inequality threatens global extreme poverty reduction effortsCoverage 14 3. 2017-10-23Brazil Poverty Rises as Social Program Coverage DropsCoverage 17 4. 2019-03-26Rahul Gandhi Promises Surgical Strike on Poverty via Income GuaranteeCoverage 101 5. 2020-11-26Global Poverty Reduction Innovation Seminar Held in Beijing Amid PandemicCoverage 14 6. 2021-02-08CBO: $15 Minimum Wage Cuts Poverty But Kills 1.4 Million JobsCoverage 551 7. 2022-02-22US Fails to Renew Child Tax Credit Despite Poverty Reduction SuccessCoverage 254 8. 2024-09-26Argentina Poverty Spikes to 53% Under Milei's Austerity Shock TherapyCoverage 145 9. 2025-11-01Kerala Becomes First Indian State To Eradicate Extreme PovertyCoverage 296 10. 2026-04-14UNDP Warns West Asia Conflict Risks Pushing 2.5 Million Indians Into PovertyCoverage 457 Browse categories to compare definitions, World V1.2 statistics and representative classified events. Example usage 1 of 3 ### Example Usage: Finance-target share doubles after COP29 as reporting shifts toward funding mechanisms and implementation The comparison isolates finance targets inside broad COP29 reporting rather than preselecting finance language. Their share rises from 21.0% during the summit to 40.7% over the following 28 days, a 19.7-point increase. The shift shows that post-summit attention moves toward funding mechanisms even as overall conference coverage decays. [[ 3 ]](https://unfccc.int/process-and-meetings/the-paris-agreement/the-glasgow-climate-pact-key-outcomes-from-cop26) [[ 4 ]](https://unfccc.int/news/cop27-reaches-breakthrough-agreement-on-new-loss-and-damage-fund-for-vulnerable-countries) [[ 5 ]](https://unfccc.int/cop28/5-key-takeaways) Earlier summits show a median post-event decline of 8.1 points, making COP29 directionally different. The measure combines climate-finance target 13.a and development-finance targets 17.1–17.5. It tracks reporting composition, not funded commitments, disbursements, emissions outcomes, policy effectiveness or any verified post-summit market repricing across exposed assets during the following month for investors. [[ 6 ]](https://unfccc.int/news/cop29-un-climate-conference-agrees-to-triple-finance-to-developing-countries-protecting-lives-and) [[ 7 ]](https://www.ipcc.ch/report/ar6/wg3/chapter/chapter-15/) COP29 summit **21.0%** 28 days after **40.7%** Composition change **+19.7pp** Prior median **-8.1pp** **COP29 reverses the prior-summit median decline** after share minus summit share Positive change Negative change Faded bars · fewer than 30 events The cohort requires an explicit COP or United Nations climate-conference reference. Each summit uses official dates plus 28 days before and after. Finance targets combine climate-finance target 13.a with development-finance targets 17.1-17.5. COP29 is compared with earlier summits, not treated as a general COP effect. The measure tracks reporting, not committed capital, disbursed finance, emissions reductions or investment returns. Example usage 2 of 3 ### Example Usage: US climate policy shifts from broad ambition toward clean-energy implementation across federal reporting The Inflation Reduction Act changes the goal composition of United States climate-policy reporting after enactment on 16 August 2022. Affordable and Clean Energy rises from 21.9% to 43.3% of assigned events, while Climate Action falls from 50.0% to 30.0%. Reporting moves from broad climate ambition toward energy implementation channels. [[ 8 ]](https://www.energy.gov/edf/inflation-reduction-act-2022) The distinction matters for capital-market research because the goals imply different exposure maps. Climate Action captures broad policy and resilience. Affordable and Clean Energy points toward power generation, efficiency and infrastructure. SDG classifications let researchers separate policy announcement from investable implementation without assuming that classified attention equals delivered spending, capacity or emissions reduction. [[ 1 ]](https://digitallibrary.un.org/record/3923923?ln=en) Affordable and Clean Energy +21.4pp Climate Action -20.0pp US climate policy shifts from broad ambition toward clean-energy implementation across federal reporting Before enactment After enactment 21.9% 43.3% **Affordable and Clean Energy** 50.0% 30.0% **Climate Action** The fixed English cohort names the Inflation Reduction Act with climate, energy or emissions language around 16 August 2022. Within-phase percentages describe assigned Sustainable Development Goal composition during the matched policy window. Example usage 3 of 3 ### Example Usage: Europe's gas crisis shifts from energy supply toward winter demand reduction Europe's gas crisis sharply increases normalized Sustainable Development Goal coverage after Russia invades Ukraine. Weekly assigned attention peaks at 1,589 events per million, more than four times the prewar median. The pooled rate rises from 335 before the war to 644 through winter, showing persistent development and investment concern. [[ 9 ]](https://www.iea.org/reports/world-energy-outlook-2022/executive-summary) The implementation channel changes through winter. Affordable and Clean Energy falls from 68.5% before the war to 54.0%, while Responsible Consumption and Production rises from 6.5% to 18.8%. The categories reveal a shift from securing supply toward conservation, efficiency and demand reduction that a single energy keyword cannot separate. [[ 10 ]](https://www.iea.org/commentaries/europes-energy-crisis-what-factors-drove-the-record-fall-in-natural-gas-demand-in-2022) Normalized coverage +309.2 per million Affordable and Clean Energy -14.5pp Responsible Consumption +12.3pp Eight-week pooled share of SDG-assigned gas-crisis events · % Responsible Consumption Affordable and Clean Energy The fixed English cohort covers European gas-supply and natural-gas reporting from September 2021 through March 2023. Each line pools goal counts across the current and prior seven weeks, then divides by all SDG-assigned cohort events in that window. Phase summaries compare the prewar period with the following winter demand-reduction phase across the same European cohort through winter 2023. Data and sources ## Download UN SDGs and References ### Downloads Download categories and counts as CSV, or the complete machine-readable release as JSON. [**CSV** Definitions and counts Download ↓](https://nosible.com/data/world-v1.2/ontologies/sustainable-development-goals/v2/nodes.csv) [**JSON** Full machine-readable release Download ↓](https://nosible.com/data/world-v1.2/ontologies/sustainable-development-goals/v2/statistics.json) Release details and citation files [Manifest](https://nosible.com/data/world-v1.2/ontologies/sustainable-development-goals/v2/manifest.json) [Release README](https://nosible.com/data/world-v1.2/ontologies/sustainable-development-goals/v2/README.md) [CC0 license and scope](https://nosible.com/data/world-v1.2/ontologies/sustainable-development-goals/v2/LICENSE.md) [Citation file](https://nosible.com/data/world-v1.2/ontologies/sustainable-development-goals/v2/CITATION.cff) [Changelog](https://nosible.com/data/world-v1.2/ontologies/sustainable-development-goals/v2/CHANGELOG.md) Need the complete World event schema? [Open the World data dictionary.](https://nosible.com/data-dictionaries#world) ### References 1. [ 1 ][United Nations General Assembly. (2015). Transforming our world: the 2030 Agenda for Sustainable Development. A/RES/70/1.](https://digitallibrary.un.org/record/3923923?ln=en) 2. [ 2 ][United Nations Department of Economic and Social Affairs. Goal 13: Take urgent action to combat climate change and its impacts.](https://sdgs.un.org/goals/goal13) 3. [ 3 ][United Nations Framework Convention on Climate Change. The Glasgow Climate Pact: key outcomes from COP26.](https://unfccc.int/process-and-meetings/the-paris-agreement/the-glasgow-climate-pact-key-outcomes-from-cop26) 4. [ 4 ][United Nations Framework Convention on Climate Change. (2022). COP27 reaches breakthrough agreement on a loss and damage fund.](https://unfccc.int/news/cop27-reaches-breakthrough-agreement-on-new-loss-and-damage-fund-for-vulnerable-countries) 5. [ 5 ][United Nations Framework Convention on Climate Change. COP28: what was achieved and what happens next?](https://unfccc.int/cop28/5-key-takeaways) 6. [ 6 ][United Nations Framework Convention on Climate Change. (2024). COP29 agrees a new collective quantified goal on climate finance.](https://unfccc.int/news/cop29-un-climate-conference-agrees-to-triple-finance-to-developing-countries-protecting-lives-and) 7. [ 7 ][Intergovernmental Panel on Climate Change. (2022). Climate Change 2022: Mitigation of Climate Change, Chapter 15: Investment and finance.](https://www.ipcc.ch/report/ar6/wg3/chapter/chapter-15/) 8. [ 8 ][U.S. Department of Energy. Inflation Reduction Act of 2022.](https://www.energy.gov/edf/inflation-reduction-act-2022) 9. [ 9 ][International Energy Agency. (2022). World Energy Outlook 2022: executive summary.](https://www.iea.org/reports/world-energy-outlook-2022/executive-summary) 10. [ 10 ][International Energy Agency. (2023). Europe's energy crisis: factors behind the record fall in natural gas demand in 2022.](https://www.iea.org/commentaries/europes-energy-crisis-what-factors-drove-the-record-fall-in-natural-gas-demand-in-2022) Continue exploring ## Complementary ontologies [43 categories EM-DAT Disaster Classification Connect development targets to the physical disaster categories shaping climate adaptation and finance. Explore ontology →](https://nosible.com/ontologies/em-dat) [300 categories ICD-11 Chapters and Blocks Link health-related development targets to clinically defined event cohorts and changing disease attention. Explore ontology →](https://nosible.com/ontologies/icd-11) > Explore 186 UN Sustainable Development Goal categories and example usages on climate finance, clean-energy implementation and European gas-demand reduction. **URL:** https://nosible.com/ontologies/sustainable-development-goals --- --- title: "Data dictionary index" description: "Field-level data dictionaries for the NOSIBLE World event dataset and every NOSIBLE Search API endpoint, each available as a PDF download." url: "https://nosible.com/data-dictionaries" --- [Home](https://nosible.com/) Data dictionaries NOSIBLE Research / Reference # Data dictionary index One data dictionary per NOSIBLE data product. Each document defines every field in its payload; endpoint dictionaries also cover request headers, error responses, and end with a worked end-to-end example. - 12 documents - Format: PDF - Updated 2026-07-20 Jump to [01 NOSIBLE World Event dataset · V1.2 · 1 document ↓](https://nosible.com/data-dictionaries#world) [02 NOSIBLE Search Search API · V2.1.0 · 11 documents ↓](https://nosible.com/data-dictionaries#search) ## 01 NOSIBLE World Event dataset · V1.2 · as of 2026-07-13 The point-in-time world event dataset. One document defines the full event payload, delivered as events.ndjson. | Document | Covers | PDF | | --- | --- | --- | | [NOSIBLE World](https://nosible.com/files/data-dictionaries/NOSIBLE-World-V1.2-Data-Dictionary.pdf) Every field in a World event, in payload order. One event per line of events.ndjson. | events.ndjson | [Download ↓ 133 KB](https://nosible.com/files/data-dictionaries/NOSIBLE-World-V1.2-Data-Dictionary.pdf) | Need category definitions and classification evidence? [Explore the World ontology reference.](https://nosible.com/ontologies) ## 02 NOSIBLE Search Search API · V2.1.0 · as of 2026-07-20 The NOSIBLE Search API. One document per endpoint; groups and order mirror the live API reference. | Document | Covers | PDF | | --- | --- | --- | | Web search Find relevant content published on the web. | | | | [Fast Search](https://nosible.com/files/data-dictionaries/NOSIBLE-Fast-Search-Data-Dictionary-V2.1.0.pdf) Synchronous search returning up to 100 results per query. | POST /search/v2/fast-search | [Download ↓ 160 KB](https://nosible.com/files/data-dictionaries/NOSIBLE-Fast-Search-Data-Dictionary-V2.1.0.pdf) | | [Time Search](https://nosible.com/files/data-dictionaries/NOSIBLE-Time-Search-Data-Dictionary-V2.1.0.pdf) One independent search per interval across a requested period. Asynchronous. | POST /search/v2/time-search | [Download ↓ 195 KB](https://nosible.com/files/data-dictionaries/NOSIBLE-Time-Search-Data-Dictionary-V2.1.0.pdf) | | [Rich Search](https://nosible.com/files/data-dictionaries/NOSIBLE-Rich-Search-Data-Dictionary-V2.1.0.pdf) Every Fast Search input plus enrichment toggles: embeddings, quant signals, classifications, website insights. | POST /search/v2/rich-search | [Download ↓ 213 KB](https://nosible.com/files/data-dictionaries/NOSIBLE-Rich-Search-Data-Dictionary-V2.1.0.pdf) | | [Bulk Search](https://nosible.com/files/data-dictionaries/NOSIBLE-Bulk-Search-Data-Dictionary-V2.1.0.pdf) Asynchronous. Up to 10,000 results per query, delivered as an encrypted file on S3. | POST /search/v2/bulk-search | [Download ↓ 189 KB](https://nosible.com/files/data-dictionaries/NOSIBLE-Bulk-Search-Data-Dictionary-V2.1.0.pdf) | | Web agents Prompt your way to perfect search results. | | | | [AI Search](https://nosible.com/files/data-dictionaries/NOSIBLE-AI-Search-Data-Dictionary-V2.1.0.pdf) Agent-planned search: builds a query strategy, probes, expands, runs a deep 100-result search. | POST /search/v2/search | [Download ↓ 99 KB](https://nosible.com/files/data-dictionaries/NOSIBLE-AI-Search-Data-Dictionary-V2.1.0.pdf) | | Web scraper Turn messy HTML pages into beautiful JSON. | | | | [Scrape URL](https://nosible.com/files/data-dictionaries/NOSIBLE-Scrape-URL-Data-Dictionary-V2.1.0.pdf) One URL in, one structured document out: text, metadata, authors, dates, snippets. | POST /search/v2/scrape-url | [Download ↓ 132 KB](https://nosible.com/files/data-dictionaries/NOSIBLE-Scrape-URL-Data-Dictionary-V2.1.0.pdf) | | Web mentions Turn keyword queries into mention time-series. | | | | [Topic Trend](https://nosible.com/files/data-dictionaries/NOSIBLE-Topic-Trend-Data-Dictionary-V2.1.0.pdf) Keyword query in, mention time-series out. | POST /search/v2/topic-trend | [Download ↓ 83 KB](https://nosible.com/files/data-dictionaries/NOSIBLE-Topic-Trend-Data-Dictionary-V2.1.0.pdf) | | Web alerts [beta] Monitor the web for exciting new content. | | | | [Save Search](https://nosible.com/files/data-dictionaries/NOSIBLE-Save-Search-Data-Dictionary-V2.1.0.pdf) Saves a search configuration to the crawl feed, which runs it on an ongoing basis. | POST /search/v2/save-search | [Download ↓ 165 KB](https://nosible.com/files/data-dictionaries/NOSIBLE-Save-Search-Data-Dictionary-V2.1.0.pdf) | | [Delete Search](https://nosible.com/files/data-dictionaries/NOSIBLE-Delete-Search-Data-Dictionary-V2.1.0.pdf) Permanently deletes a saved search and turns off its web alert. | POST /search/v2/delete-search | [Download ↓ 72 KB](https://nosible.com/files/data-dictionaries/NOSIBLE-Delete-Search-Data-Dictionary-V2.1.0.pdf) | | [Get Searches](https://nosible.com/files/data-dictionaries/NOSIBLE-Get-Searches-Data-Dictionary-V2.1.0.pdf) Returns every saved search on your API key, with full configurations. | POST /search/v2/get-searches | [Download ↓ 79 KB](https://nosible.com/files/data-dictionaries/NOSIBLE-Get-Searches-Data-Dictionary-V2.1.0.pdf) | Endpoint dictionaries are versioned to the Search API release published in the [canonical API docs](https://docs.nosible.com/). Field names, types, shapes, and example values are documented per field. [Explore the Search API.](https://docs.nosible.com/) > Field-level data dictionaries for the NOSIBLE World event dataset and every NOSIBLE Search API endpoint, each available as a PDF download. **URL:** https://nosible.com/data-dictionaries --- --- title: "Compliance and trust" description: "Review NOSIBLE's policies for public-web crawling, copyright, privacy, and security before procurement, legal review, or deployment." url: "https://nosible.com/legal/compliance" --- Trust / Buyer diligence Updated 2026-07-21 # Compliance and trust Four policies explain how NOSIBLE collects, handles, and protects data. Start with crawling to understand what enters the index. Then review how NOSIBLE handles copyrighted content, personal data, and customer access. Each page states the operating boundary, supporting evidence, and answers to common diligence questions. NOSIBLE is operated by Nosible Inc. Four review areas ## Open the policy that answers your diligence question [01 Public collection Crawling What enters the index? NOSIBLE crawls public written content. It does not bypass paywalls, logins, or other access controls. Source and domain review Robots.txt and crawl controls Point-in-time date verification Review policy →](https://nosible.com/legal/crawling) [02 Content rights Copyright What leaves the index? Search results return limited snippets, metadata, attribution, and source links—not cached articles or full-text resale. Non-substitutive search Snippet limits and attribution Jurisdiction-specific legal review Review policy →](https://nosible.com/legal/copyright) [03 Personal data Privacy What personal data is retained? NOSIBLE minimizes personal data, limits operational retention, and provides clear removal, objection, and delisting routes. Index exclusions API data minimization Removal and delisting requests Review policy →](https://nosible.com/legal/privacy) [04 Customer access Security How is customer access protected? NOSIBLE documents the infrastructure, access controls, monitoring, backups, and deployment options protecting customer workflows. Infrastructure and access Backups and replication Operational security controls Review policy →](https://nosible.com/legal/security) > Review NOSIBLE's policies for public-web crawling, copyright, privacy, and security before procurement, legal review, or deployment. **URL:** https://nosible.com/legal/compliance --- --- title: "Our posture on crawling" description: "How NOSIBLE crawls public written content, excludes high-risk sources, prioritizes URLs, and verifies point-in-time dates." url: "https://nosible.com/legal/crawling" --- Crawling # Our posture on crawling NOSIBLE crawls and indexes publicly available written content and cross-references it so customers can find what was verifiably public knowledge at any point in time. We do not crawl behind paywalls or logins. We follow ethical crawling principles. This puts NOSIBLE on the public-access side of [hiQ v. LinkedIn](https://cdn.ca9.uscourts.gov/datastore/opinions/2022/04/18/17-16783.pdf), [Van Buren](https://www.supremecourt.gov/opinions/20pdf/19-783_k53l.pdf), and related U.S. authority. NOSIBLE is operated by Nosible Inc. - NOSIBLE crawls public written content. No paywalls, logins, or bypass. - The index has 2.5B+ pages from 300,000+ sources in 150+ countries. - Coverage spans 95 languages across news, company, government, and specialist sources. - Every crawled domain must trace to a legal entity and pass website checks. - Website checks cover redirects, blocklists, Web Risk, RDAP, robots.txt, and sitemaps. - NOSIBLE honors robots.txt and avoids the conduct in [hiQ](https://cdn.ca9.uscourts.gov/datastore/opinions/2022/04/18/17-16783.pdf), [Power Ventures](https://cdn.ca9.uscourts.gov/datastore/opinions/2016/07/12/13-17102.pdf), and [3Taps](https://law.justia.com/cases/federal/district-courts/FSupp2/942/962/). - Fresh URLs get priority. We always enforce a reasonable crawl rate across all approved websites. - NOSIBLE supports ETags and Last-Modified headers to minimize unnecessary load on host servers. - Proxy networks we work with must source IPs legally and ethically at scale. - Published dates are verified before entering research use or model workflows. ## NOSIBLE crawls public written sources for search and retrieval NOSIBLE is focused on long-form written content from news websites, press releases, company websites, specialist publications, government pages, public records, official filings, court documents, and other vetted public sources. The product is built for discovery, retrieval, source ranking, and point-in-time verification. It is not a cached-page reader, publisher archive, marketplace mirror, social archive, or paywall bypass. U.S. authority draws a strong distinction between public, logged-out access and attempts to enter gated systems, evade blocks, or continue after individualized revocation ( [Van Buren](https://www.supremecourt.gov/opinions/20pdf/19-783_k53l.pdf); [hiQ](https://cdn.ca9.uscourts.gov/datastore/opinions/2022/04/18/17-16783.pdf); [Power Ventures](https://cdn.ca9.uscourts.gov/datastore/opinions/2016/07/12/13-17102.pdf); [Craigslist v. 3Taps](https://law.justia.com/cases/federal/district-courts/FSupp2/942/962/) ). NOSIBLE stays in the public-page category: responsible crawling, no access-control bypass, and links back to the original source. Public, logged-out scraping also has support on contract-law grounds. [Meta v. Bright Data](https://www.courthousenews.com/wp-content/uploads/2024/01/meta-platforms-v-bright-data-ruling-motion-for-summary-judgment.pdf) rejected Meta's contract claims against logged-out public scraping, while [Nguyen v. Barnes & Noble](https://cdn.ca9.uscourts.gov/datastore/opinions/2014/08/18/12-56628.pdf) shows why browsewrap terms without affirmative assent are weaker than click-to-accept terms. ## Crawl scope is defined before a page enters the index | Area | NOSIBLE position | | --- | --- | | Allowed source types | News, press releases, company pages, specialist sources, government pages, public records, official filings, and court documents with useful written evidence. | | Excluded source types | Paywalls, logged-in pages, disallowed social platforms, marketplaces, unsafe domains screened with [Google Web Risk](https://cloud.google.com/security/products/web-risk), inherently pornographic sites, and pages containing personally identifiable information. | | Domain review | Every crawled domain must be traceable to a legal entity and pass website checks before it enters the source universe or production index. | | Website verification | Before approval, NOSIBLE records URL preflight results, homepage response metadata, redirect status, blocklist status, Google index presence, Google Web Risk status, RDAP registration signals, robots.txt, ads.txt, security.txt, legal-document paths, content sitemaps, and archive provenance. | | URL priority | Fresh URLs are prioritized over historical URLs, with reasonable crawl rates maintained across source websites. | | Crawl impact controls | Robots.txt, no access-control bypass, no block evasion, ETag and Last-Modified checks, cache-aware requests, and responsible crawl rates. | | Point-in-time checks | Published dates are checked against metadata, modified dates, temporal falsification tests, corroborating sources, systems of record, point-in-time web archives such as Common Crawl, and sensibility checks. | ## Published dates are verified before they enter research workflows A published date is not accepted as ground truth. NOSIBLE treats it as a claim that must be supported or rejected before a backtest or AI agent can use the source. Customers can use stricter evidence thresholds when confidence matters more than volume. Public-source personal data still carries privacy obligations under ICO web-scraping guidance and UK/EU GDPR ( [ICO](https://ico.org.uk/about-the-ico/what-we-do/our-work-on-artificial-intelligence/response-to-the-consultation-series-on-generative-ai/the-lawful-basis-for-web-scraping-to-train-generative-ai-models/); [GDPR Article 6](https://gdpr-info.eu/art-6-gdpr/) ). ## Common crawling questions ### Does NOSIBLE crawl behind paywalls? No. NOSIBLE does not crawl behind paywalls, logged-in areas, subscription gates, or other non-public access controls. U.S. authority treats public, logged-out web access differently from entering gated systems or bypassing access controls ( [Van Buren](https://www.supremecourt.gov/opinions/20pdf/19-783_k53l.pdf); [hiQ](https://cdn.ca9.uscourts.gov/datastore/opinions/2022/04/18/17-16783.pdf) ). NOSIBLE operates on the public side of that line. ### Does NOSIBLE crawl social media or marketplaces? Partially. NOSIBLE considers websites case by case and has recently begun crawling some public, logged-out LinkedIn pages. NOSIBLE does not currently crawl Facebook, Instagram, X (formerly Twitter), Reddit, Amazon, eBay, Airbnb, or similar platforms. It does not crawl behind paywalls, logged-in areas, subscription gates, or other non-public access controls. Social platforms receive additional review because source sensitivity, anti-scraping signals, and access controls matter to a lawful web-scraping assessment ( [CNIL](https://www.cnil.fr/en/legal-basis-legitimate-interest-focus-sheet-measures-implement-case-data-collection-web-scraping) ). ### Does NOSIBLE snapshot web pages? No. NOSIBLE focuses on crawling evergreen content pages for search and retrieval, not visual snapshots or cached-page browsing. The product stores technical crawl artifacts so relevant passages can be found later, but it is not designed to let users browse historical page copies. ### What source types does NOSIBLE crawl? NOSIBLE crawls news stories, press releases, company websites, specialist publications, government websites, public records, official filings, court documents, niche financial publications, non-financial specialist sites, blogs, and certain high-quality forums. The common thread is that the content is public, written, source-like, and useful for search or verification. ### What source types does NOSIBLE exclude before indexing? NOSIBLE excludes websites flagged by internal blocklists or [Google Web Risk](https://cloud.google.com/security/products/web-risk), inherently pornographic websites, disallowed social platforms, marketplace and listing websites, paywalled content, and pages containing personally identifiable information. Website checks also cover redirects, RDAP registration signals, robots.txt, and sitemaps before a source is approved. The crawler stays focused on pages that improve search and avoids categories that add risk without improving the core evidence layer. ### How large is the crawl and index? NOSIBLE's index contains more than 2.5 billion pages from over 300,000 sources in more than 150 countries and 95 languages. Point-in-time verification, domain authority, source comparison, and cross-source corroboration all improve when the search engine has more independent anchors. ### How much first-hand history does NOSIBLE have? NOSIBLE has first-hand point-in-time web data going back to 2023. Earlier history is derived through evidence checks, corroborating sources, systems of record, and archive-based verification. First-hand crawls and reconstructed historical evidence carry different confidence levels, so customers can apply stricter thresholds where the workflow requires them. ### How often does NOSIBLE check whitelisted sources? NOSIBLE checks vetted sources frequently so new public pages can enter future crawl batches quickly. The crawler prioritizes freshness while maintaining reasonable crawl rates. Domains must be traceable to a legal entity and pass website checks before they enter the source universe. ### How does NOSIBLE verify a website before indexing? NOSIBLE starts with a URL preflight check, then records the normalized scheme, host, domain, suffix, geography, homepage response, redirect behavior, status code, HTTP version, elapsed time, title, and language. The domain is checked against internal blocklists, Google Web Risk, Google index presence, RDAP registration records, robots.txt, content sitemaps, ads.txt, security.txt, privacy policy links, terms links, and archive provenance before it becomes a production source. ### How does NOSIBLE identify the legal entity behind a website? NOSIBLE looks for ownership clues in the homepage footer, copyright strings, privacy policies, and terms of service. It extracts the legal entity name and jurisdiction where the evidence is available. RDAP registration data, ads.txt owner-domain fields, and public legal documents give additional signals about who operates the site. ### How does NOSIBLE use archive provenance for website verification? NOSIBLE checks domain history through point-in-time web archives such as Common Crawl where available. Those records help estimate when a domain first appeared, when it was last seen, and whether its history is consistent with the source being evaluated. ### How does NOSIBLE decide which URLs to crawl first? Every domain NOSIBLE crawls must be traceable to a legal entity and pass website checks. Fresh URLs are prioritized over historical URLs, and NOSIBLE maintains reasonable crawl rates across source websites. The crawler also looks for ETags and Last-Modified headers so repeat checks can use conditional requests where the server supports them. ### How does NOSIBLE reduce the burden on source websites? NOSIBLE uses several crawl-impact controls. It looks for ETags and Last-Modified values, then uses conditional requests where the source server supports them. It also maintains reasonable crawl rates. Trespass theories require actual server damage or impairment, not mere dislike of crawling ( [Intel v. Hamidi](https://scocal.stanford.edu/opinion/intel-v-hamadi-33288) ); [eBay v. Bidder's Edge](https://law.justia.com/cases/federal/district-courts/FSupp2/100/1058/2478126/) remains a high-volume, fact-bound warning case. ### Does NOSIBLE use proxies or country-aware routing? Yes. NOSIBLE owns thousands of dedicated IP addresses and tries to access pages from the country where the website operates before using more advanced proxies. NOSIBLE depends on Bright Data for most crawling efforts and has fallback providers in place. Proxies are for reliable public-page access and routing, not for evading individualized blocks, access controls, or revocation notices ( [3Taps](https://law.justia.com/cases/federal/district-courts/FSupp2/942/962/) ). ### Does NOSIBLE use web archives? Yes. NOSIBLE uses web archives to discover URLs for recently whitelisted domains and adds those URLs to the backlog. NOSIBLE also ingests the HTML of missing pages that return 404 Not Found directly from web archives. Content sourced directly from web archives is tagged so its provenance is transparent and available for filtering. ### Can customers influence which sites NOSIBLE indexes? Yes. Customers can suggest sites for NOSIBLE to add, prioritize, or deprioritize when those sources matter to a research or risk workflow. Suggestions do not change crawl scope. A requested source still needs to be public, written, reachable without access-control bypass, and consistent with the exclusion policy before it enters the index. ### Does NOSIBLE wrap Brave, Exa, or another search engine? No. NOSIBLE does not wrap Brave, Exa, or any other search API. The index is independent and built from NOSIBLE's own crawling, parsing, routing, and retrieval systems. Customers get NOSIBLE source coverage, timing controls, and retrieval logic rather than a repackaged third-party search feed. ### Does NOSIBLE advertise a differentiated user agent? No. Like [Brave Search](https://search.brave.com/help/brave-search-crawler), NOSIBLE does not advertise a differentiated crawler user agent because some websites selectively allow only a small set of recognized search crawlers. This does not change NOSIBLE's access rules: it crawls only public pages, honors robots.txt, does not bypass access controls, and does not crawl a page that is unavailable to Googlebot. ### How does NOSIBLE verify point-in-time dates? NOSIBLE treats a published date as a claim, not a fact. The process checks HTML metadata, structured JSON, modified timestamps, temporal falsification tests, corroborating sources, authoritative systems of record such as [EDGAR](https://www.sec.gov/edgar), point-in-time web archives such as Common Crawl, and basic sensibility checks. These checks reduce future information leaking into historical research, backtests, and agent workflows. ### What does temporal falsification mean? Temporal falsification means checking whether the page contains information that should not exist at the claimed date. Examples include references to ChatGPT before 2022, COVID-19 before 2020, Oumuamua before 2017, Ozempic before 2008, or YouTube before 2005. If a page makes those kinds of impossible references, the claimed date needs to be rejected or treated with much lower confidence. ### Can customers choose stricter point-in-time evidence thresholds? Yes. Different firms may require different amounts of evidence before accepting that a page was available at a claimed point in time. NOSIBLE can dial the evidence threshold up or down. A stricter threshold improves confidence and reduces look-ahead risk, but it also reduces volume because fewer URLs will clear the bar. ### Why does NOSIBLE store technical crawl artifacts? NOSIBLE retains a slimmed-down, minified, compressed version of HTML for technical purposes. It also stores structured JSON that converts raw HTML into human-readable text and chunks. Those chunks enter the search index so NOSIBLE can retrieve relevant passages in response to customer queries. Users get retrieval infrastructure, not a full-page reading product. ### What website-level metadata does NOSIBLE maintain? NOSIBLE maintains domain-level metadata including host, title, description, language, geography, industry, topic, brand-safety status, indexed totals, first published date, last published date, sitemap status, legal-document paths, RDAP signals, archive provenance, shard routing, and content trend. The system uses this metadata to understand source quality before a query is run. ### How does NOSIBLE use domain authority? NOSIBLE uses website metadata to compute query-specific domain-authority scores. Those scores help rank sources when many pages match the same topic. Authority scoring can improve signal quality in workflows such as sentiment analysis, risk monitoring, and source comparison. Other legal and compliance policies [Compliance index](https://nosible.com/legal/compliance) [Copyright](https://nosible.com/legal/copyright) [Privacy](https://nosible.com/legal/privacy) [Security](https://nosible.com/legal/security) [Start Trial](https://nosible.com/start-trial) [Return to Trial Checklist](https://nosible.com/start-trial#legal) > How NOSIBLE crawls public written content, excludes high-risk sources, prioritizes URLs, and verifies point-in-time dates. **URL:** https://nosible.com/legal/crawling --- --- title: "Our posture on copyright" description: "NOSIBLE's copyright position for search, snippets, attribution, non-substitution, technical copying, and legal review." url: "https://nosible.com/legal/copyright" --- Copyright # Our posture on copyright NOSIBLE uses public pages to build search, not to republish articles. The copyright question is whether the product helps users find and evaluate sources or replaces the original work. NOSIBLE is built for discovery, attribution, and non-substitutive retrieval ( [Authors Guild v. Google](https://law.justia.com/cases/federal/appellate-courts/ca2/13-4829/13-4829-2015-10-16.html) ). NOSIBLE is operated by Nosible Inc. - Facts are not copyrightable. Hot-news claims remain narrow under [Feist](https://supreme.justia.com/cases/federal/us/499/340/) and [NBA v. Motorola](https://law.justia.com/cases/federal/appellate-courts/F3/105/841/598844/). - NOSIBLE is not trying to keep users away from source websites. It is trying to help them find the right source faster. - NOSIBLE builds search, rankings, metadata, timestamps, and snippets from public pages. - Users get short snippets, source metadata, dates, and links back to originals. - NOSIBLE is for discovery and evidence review, not article consumption. - It does not sell cached pages, article archives, full text, or publisher replacement. - Rate limits and abuse monitoring reduce full-work reconstruction risk. - NOSIBLE respects robots.txt, rights reservations, and source-removal requests. - Search cases support indexing for discovery: [Authors Guild](https://law.justia.com/cases/federal/appellate-courts/ca2/13-4829/13-4829-2015-10-16.html), [HathiTrust](https://law.justia.com/cases/federal/appellate-courts/ca2/12-4547/12-4547-2014-06-10.html), [Kelly](https://law.justia.com/cases/federal/appellate-courts/F3/336/811/468703/), and [Perfect 10](https://law.justia.com/cases/federal/appellate-courts/F3/508/1146/). ## NOSIBLE uses public pages to build search, not to republish articles Copyright law distinguishes technical copying for discovery from redistribution for consumption. Search and snippet cases protect transformative, limited, non-substitutive uses, including [Authors Guild v. Google](https://law.justia.com/cases/federal/appellate-courts/ca2/13-4829/13-4829-2015-10-16.html), [HathiTrust](https://law.justia.com/cases/federal/appellate-courts/ca2/12-4547/12-4547-2014-06-10.html), [Kelly v. Arriba Soft](https://law.justia.com/cases/federal/appellate-courts/F3/336/811/468703/), [Perfect 10 v. Amazon](https://law.justia.com/cases/federal/appellate-courts/F3/508/1146/), and [A.V. v. iParadigms](https://www.copyright.gov/fair-use/summaries/a.v.-vanderhye-iparadigms-4thcir2009.pdf). NOSIBLE uses public pages for discovery, attribution, ranking, and source retrieval. [Google v. Oracle](https://www.supremecourt.gov/opinions/20pdf/18-956_d18f.pdf) confirms that commercial use does not defeat fair use where the use adds a different function and does not substitute for the original market. [Warhol v. Goldsmith](https://www.supremecourt.gov/opinions/22pdf/21-869_87ad.pdf) reinforces the need to keep purpose different from the original work. NOSIBLE's purpose is search and source discovery, not article consumption. ## Copyright risk is controlled by transformation, limits, and attribution | Area | NOSIBLE position | | --- | --- | | Transformative use | Public pages become indexes, entity signals, timestamps, source metadata, shards, relevance scores, and ranked snippets. | | Snippet limits | Returned text is limited to what is needed to judge relevance. It should answer whether the source matters, not replace the work. | | Attribution | Every result identifies the source and routes the user back to the original page. | | Non-substitution | NOSIBLE is designed for source discovery, not article consumption, archive reading, or full-text resale. | | Opt-out handling | NOSIBLE respects robots.txt and machine-readable reservations, which supports implied-license and lawful-access arguments ( [Field v. Google](https://www.law.berkeley.edu/archive/files/Field_v_Google.pdf); [EU Directive 2019/790](https://eur-lex.europa.eu/eli/dir/2019/790/oj/eng) ). | | Anti-extraction | Retrieval endpoints are rate-limited and monitored so customers cannot reconstruct complete works at scale. | | Facts and events | Facts are not copyrightable, and hot-news claims remain narrow ( [Feist](https://supreme.justia.com/cases/federal/us/499/340/); [NBA v. Motorola](https://law.justia.com/cases/federal/appellate-courts/F3/105/841/598844/) ). | | Legal review | Enterprise terms and indemnity can be discussed based on the use case, deployment model, contract, and risk allocation. | ## Snippets are better for AI systems and safer for publishers Short snippets are a better technical unit for search, AI, and quantitative analysis. Full text often includes navigation, boilerplate, duplicate paragraphs, unrelated context, and long passages that dilute the signal. A snippet gives the model or analyst the relevant evidence and leaves the original source in control of the full work. Search indexing, snippet view, temporary technical copying, short extracts, hyperlinking, and structured fairness analysis are supported across major jurisdictions, including the U.S. search cases above, [EU Directive 2019/790](https://eur-lex.europa.eu/eli/dir/2019/790/oj/eng), [EU Directive 2001/29/EC](https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32001L0029), [UK CDPA section 28A](https://www.legislation.gov.uk/ukpga/1988/48/section/28A), and [CCH Canadian Ltd. v. Law Society of Upper Canada](https://scc-csc.lexum.com/scc-csc/scc-csc/en/item/2125/index.do). ## Common copyright questions ### Does NOSIBLE republish articles? No. NOSIBLE does not sell a reading product, archive product, cached-page product, or full-text resale product. It returns relevance-bound snippets, metadata, dates, and links back to the original source. Search cases protect tools that help users find and evaluate works without giving them a substitute for the works ( [Authors Guild v. Google](https://law.justia.com/cases/federal/appellate-courts/ca2/13-4829/13-4829-2015-10-16.html) ). ### What is NOSIBLE's copyright basis? NOSIBLE is non-substitutive search. It transforms public pages into indexes, rankings, metadata, shards, relevance scores, and short previews so users can discover sources rather than consume original works inside NOSIBLE. [HathiTrust](https://law.justia.com/cases/federal/appellate-courts/ca2/12-4547/12-4547-2014-06-10.html) treated full-text search as different from reading the underlying work. ### Why are snippets important? Snippets show enough text to judge relevance without replacing the original work. NOSIBLE returns short snippets of 256 words or less, with source attribution. That limit matters legally because it avoids a full-text reading experience, and technically because snippets give models and analysts relevant evidence without forcing them through boilerplate, navigation, duplicate text, and unrelated context. ### Why does NOSIBLE say snippets are technically better than full text? Snippets improve signal-to-noise, normalize document length, reduce semantic dilution in vector search, fit more sources into an AI context window, reduce latency and inference cost, and reduce hallucination risk by grounding generation in specific retrieved passages. The same design that reduces substitution risk also improves retrieval quality. ### What prevents customers from reconstructing full works? NOSIBLE uses anti-extraction controls, including strict rate limits and abuse monitoring on retrieval endpoints. The product is designed around snippets, source metadata, rankings, and links, not full-text reconstruction. Fair-use analysis turns heavily on non-substitution and market effect ( [Authors Guild v. Google](https://law.justia.com/cases/federal/appellate-courts/ca2/13-4829/13-4829-2015-10-16.html); [Google v. Oracle](https://www.supremecourt.gov/opinions/20pdf/18-956_d18f.pdf) ). ### How do anti-extraction controls protect the search product? Anti-extraction protects the search design. If customers could reconstruct full works through repeated retrieval, NOSIBLE would start to look like a redistribution system instead of a search system. Rate limits and abuse monitoring help keep retrieval focused on relevance, evidence, attribution, and source discovery instead of bulk text extraction. ### How does attribution reduce copyright risk? Every result includes the source domain, original URL, and date visited. Attribution matters because it routes human and AI traffic back to the original publisher for consumption. It also reinforces the discovery function of the product. NOSIBLE is not trying to keep users away from source websites. It is trying to help them find the right source faster. ### Does NOSIBLE respect robots.txt and opt-out signals? Yes. NOSIBLE respects robots.txt and machine-readable reservations, and source owners can request removal. Robots and no-index style controls support implied-license and opt-out analysis in search cases such as [Field v. Google](https://www.law.berkeley.edu/archive/files/Field_v_Google.pdf). EU text-and-data-mining rules also recognize machine-readable rights reservations under [Directive 2019/790](https://eur-lex.europa.eu/eli/dir/2019/790/oj/eng). ### Is NOSIBLE a generative AI product that reproduces publisher text? No. NOSIBLE is search and retrieval infrastructure. It returns dated, attributed source results and snippets so users can inspect evidence and open the original source. It does not generate substitute articles or expressive outputs from publisher text. That separates NOSIBLE from disputes centered on competing expressive outputs, including [Thomson Reuters v. Ross](https://www.ded.uscourts.gov/sites/ded/files/opinions/20-613_5.pdf) and the pending OpenAI litigation. ### Why does source attribution matter for AI agents? Source attribution keeps the evidence visible. An AI agent can inspect the source domain, original URL, date visited, and relevant snippet before using a result. That reduces black-box retrieval risk and routes the user back to the original publisher when the full work needs to be read. ### How is NOSIBLE different from a web archive or scraping vendor? A web archive or bulk scraping product can create substitution risk if it gives users cached pages, full-text redistribution, or archive-style access to third-party content. NOSIBLE turns source pages into search infrastructure and returns limited snippets with attribution. Users find third-party articles in NOSIBLE and read them at the original source. ### How is NOSIBLE different from Common Crawl? NOSIBLE is a search and retrieval product, not a bulk web corpus for redistribution. It returns snippets, rankings, metadata, dates, source attribution, and links back to original pages. Common Crawl-style access can be useful for some technical teams, but NOSIBLE is built around discovery, point-in-time retrieval, and anti-extraction controls. ### Does NOSIBLE provide cached full-page access? No. NOSIBLE does not provide cached full-page access, archive-style reading, or full-text resale. The product returns snippets, metadata, rankings, dates, and links back to the source. Users discover the source in NOSIBLE and consume the full work at the original website. ### How is NOSIBLE safer than naive scraping? NOSIBLE is more conservative than products that provide cached pages, archive-style access, full-text redistribution, or bulk extraction of third-party text. It returns limited snippets with source attribution and retrieval controls. The product is built for discovery, not consumption of publisher content inside NOSIBLE. ### Which U.S. cases support the copyright position? [Authors Guild v. Google](https://law.justia.com/cases/federal/appellate-courts/ca2/13-4829/13-4829-2015-10-16.html), [HathiTrust](https://law.justia.com/cases/federal/appellate-courts/ca2/12-4547/12-4547-2014-06-10.html), [Kelly v. Arriba Soft](https://law.justia.com/cases/federal/appellate-courts/F3/336/811/468703/), [Perfect 10 v. Amazon](https://law.justia.com/cases/federal/appellate-courts/F3/508/1146/), [A.V. v. iParadigms](https://www.copyright.gov/fair-use/summaries/a.v.-vanderhye-iparadigms-4thcir2009.pdf), and [Field v. Google](https://www.law.berkeley.edu/archive/files/Field_v_Google.pdf) support the position. They protect search, indexing, snippets, thumbnails, anti-plagiarism indexing, caching, or source-location functions when the tool helps users find information without replacing the work. ### What does the EU and UK framework add? EU and UK law add jurisdiction-specific support for temporary technical copies, short extracts, hyperlinking, and text-and-data mining in defined circumstances, including the UK's temporary-copying exception in [CDPA section 28A](https://www.legislation.gov.uk/ukpga/1988/48/section/28A). The EU framework includes a commercial text-and-data-mining exception subject to rights reservations under [Directive 2019/790](https://eur-lex.europa.eu/eli/dir/2019/790/oj/eng). NOSIBLE evaluates those rules alongside its non-substitutive search design and does not treat any single exception as blanket authorization. ### What about Canada, Japan, Singapore, Israel, and Australia? The same pattern appears outside the U.S., EU, and UK. Canada treats fair dealing as a structured balancing test. Japan allows non-enjoyment uses including data analysis and machine learning. Singapore allows computational data analysis of lawfully accessed works if originals are not distributed. Israel and Australia also focus on transformation, proportionality, and market effect. ### Can enterprise customers discuss indemnity? Yes. Third-party intellectual property rights indemnity can be negotiated for enterprise customers. The terms depend on the use case, deployment model, contract, and risk allocation. Indemnity allocates commercial and legal risk between NOSIBLE and the customer. Other legal and compliance policies [Compliance index](https://nosible.com/legal/compliance) [Crawling](https://nosible.com/legal/crawling) [Privacy](https://nosible.com/legal/privacy) [Security](https://nosible.com/legal/security) [Start Trial](https://nosible.com/start-trial) [Return to Trial Checklist](https://nosible.com/start-trial#legal) > NOSIBLE's copyright position for search, snippets, attribution, non-substitution, technical copying, and legal review. **URL:** https://nosible.com/legal/copyright --- --- title: "Our posture on privacy" description: "How NOSIBLE minimizes personal data, retains limited usage logs, and handles removal, objection, erasure, and delisting requests." url: "https://nosible.com/legal/privacy" --- Privacy # Our posture on privacy NOSIBLE is built around public written content, not personal profiles. Public availability does not remove privacy rights. NOSIBLE minimizes personal data in the index, retains limited usage logs for 24 hours, does not log API request or response payloads on normal endpoints, and provides routes for objections, delisting, and removal ( [ICO](https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/individual-rights/the-right-to-be-informed/what-common-issues-might-come-up-in-practice/) ). NOSIBLE is operated by Nosible Inc. - NOSIBLE indexes public written sources, not personal profiles or identity datasets. - The search index is not designed to hold personal data or sensitive records. - The crawler excludes PII, unsafe domains, adult sites, disallowed social platforms, and marketplaces. - Website verification checks privacy and terms links for ownership, legal context, and rights routes. - Public availability does not remove privacy rights under [GDPR Article 6](https://gdpr-info.eu/art-6-gdpr/) and ICO guidance. - Usage logs retain only the operational data needed for rate limits and abuse prevention and are deleted after 24 hours. - NOSIBLE does not log or record API request or response payloads on normal endpoints. - Websites and individuals can email [stuart@nosible.com](mailto:stuart@nosible.com) to request removal, objection, erasure, or delisting review under [GDPR Article 17](https://gdpr-info.eu/art-17-gdpr/). ## Privacy starts with what does not enter the index NOSIBLE does not index personally identifiable information and does not hold personal data in the search index. The crawler also excludes unsafe domains, inherently pornographic sites, disallowed social platforms, marketplaces, and paywalled content. Privacy risk is reduced before content reaches the core product. When public-source text still contains personal data, privacy law still applies. NOSIBLE relies on legitimate interests, necessity, minimization, public transparency, and case-by-case handling of objections or removal requests, consistent with [GDPR Article 6](https://gdpr-info.eu/art-6-gdpr/), [GDPR Article 14](https://gdpr-info.eu/art-14-gdpr/), [ICO guidance](https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/lawful-basis/a-guide-to-lawful-basis/legitimate-interests/), [CNIL web-scraping guidance](https://www.cnil.fr/en/legal-basis-legitimate-interest-focus-sheet-measures-implement-case-data-collection-web-scraping), and [EDPB Opinion 28/2024](https://www.edpb.europa.eu/system/files/2024-12/edpb_opinion_202428_ai-models_en.pdf). ## Customer API data is minimized by default | Area | NOSIBLE position | | --- | --- | | Search index | The index is for public written content and source metadata. It is not designed to hold personal data. | | Website legal pages | NOSIBLE scans homepage links for privacy policies and terms of service so the source can be reviewed against public legal context. | | Legal basis | Public-source processing is assessed through legitimate interests, necessity, minimization, and balancing under [GDPR Article 6](https://gdpr-info.eu/art-6-gdpr/). | | Transparency | Web-scale indirect collection relies on public notice and safeguards under [GDPR Article 14](https://gdpr-info.eu/art-14-gdpr/) and the Article 14(5)(b) disproportionate-effort framework. | | Usage logs | Usage logs contain only the operational data needed for rate-limit enforcement and abuse prevention. They are always enabled and retained for 24 hours. | | API payloads | NOSIBLE does not log or record API request or response payloads on normal endpoints. | | Delivery files | Bulk Search and Time Search results are delivered as encrypted JSON files. | ## Source owners can ask to be removed from the index Website owners and individuals can email [stuart@nosible.com](mailto:stuart@nosible.com) with the affected URL, the request type, and enough information to verify their identity or authority. Website opt-out requests lead to removal from the index and notice to affected paying customers. Individual objections, erasure requests, and delisting requests are reviewed case by case against privacy, accuracy, source context, public role, public interest, and the right to information under [GDPR Article 17](https://gdpr-info.eu/art-17-gdpr/), [Google Spain](https://curia.europa.eu/jcms/upload/docs/application/pdf/2014-05/cp140070en.pdf), [GC and Others v. CNIL](https://curia.europa.eu/jcms/upload/docs/application/pdf/2019-09/cp190113en.pdf), [Google LLC v. CNIL](https://curia.europa.eu/jcms/upload/docs/application/pdf/2019-09/cp190112en.pdf), and [NT1 and NT2 v. Google](https://www.judiciary.uk/wp-content/uploads/2018/04/nt1-nt2-v-google-press-summary-180413.pdf). ## Common privacy questions ### Does NOSIBLE store personal data in the index? No. NOSIBLE does not index personally identifiable information and does not hold personal data in the search index. The crawler is focused on public written source material, and pages that create personal-data risk are outside the intended search corpus. Privacy risk is reduced before content reaches the product. ### Does public availability remove privacy rights? No. Public availability is relevant to reasonable expectations, but it does not switch off privacy law. ICO guidance is explicit that personal data can remain protected even when it is publicly available ( [ICO](https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/individual-rights/the-right-to-be-informed/what-common-issues-might-come-up-in-practice/) ). NOSIBLE handles that through minimization, source exclusions, public privacy information, and a route for objections, delisting, or removal requests. ### What lawful basis applies to public-source processing? For UK and EU privacy analysis, the relevant basis is legitimate interests under [GDPR Article 6(1)(f)](https://gdpr-info.eu/art-6-gdpr/). That requires a purpose, necessity, and balancing assessment. [KNLTB v. Autoriteit Persoonsgegevens](https://curia.europa.eu/juris/document/document.jsf?docid=290688&doclang=EN) confirms that commercial interests can qualify, subject to strict necessity and balancing. NOSIBLE limits processing through source context, exclusions, minimization, and rights handling. ### Why does NOSIBLE look for privacy policies and terms of service? Privacy policies and terms help identify who operates a website, what jurisdiction may apply, and what public rights routes the source owner makes available. NOSIBLE uses homepage link scanning and strict URL matching to find those documents. They support source review, legal-entity extraction, jurisdiction checks, and case-by-case handling of removal or privacy objections. ### Why is unsafe or adult content mentioned on the privacy page? Those exclusions reduce data-handling risk before content enters the product. NOSIBLE checks domains with [Google Safe Browsing](https://safebrowsing.google.com/) and excludes inherently pornographic websites using known unsafe or NSFW lists. CNIL guidance also treats sensitivity-heavy sources as higher risk in scraping assessments ( [CNIL](https://www.cnil.fr/en/legal-basis-legitimate-interest-focus-sheet-measures-implement-case-data-collection-web-scraping) ). These categories add privacy and safety risk without improving the search corpus. ### What customer data does NOSIBLE retain for API operations? NOSIBLE retains limited usage logs for 24 hours to enforce rate limits, prevent abuse, and maintain service reliability. These logs contain the operational metadata needed to run the API. NOSIBLE does not log or record search queries, request payloads, response payloads, or customer content on normal endpoints. ### Can usage logging be disabled? No. NOSIBLE keeps the minimum usage metadata required to enforce rate limits, identify abuse, and operate the service reliably. It is always enabled and automatically deleted after 24 hours. This operational logging does not include search queries, request payloads, response payloads, or customer content from normal endpoints. ### How are Bulk Search and Time Search results protected? Bulk Search and Time Search results are delivered as encrypted JSON files rather than written to ordinary API logs. Access is limited through the delivery configuration agreed with the customer. After delivery, customers control who can access the files, where they are stored, and how long downstream systems retain them. ### Can a website ask to be forgotten? Yes. A website owner can email [stuart@nosible.com](mailto:stuart@nosible.com) with the affected domain or URLs and enough information to verify ownership or authority. NOSIBLE removes approved opt-out requests from the index and notifies affected paying customers of the change. ### Can an individual ask NOSIBLE to delist a result about them? Yes. Email [stuart@nosible.com](mailto:stuart@nosible.com) with the affected result and proof of identity. NOSIBLE reviews requests case by case, balancing privacy, accuracy, public role, source context, and public interest. Delisting is not automatic; information about public figures or significant public matters may remain indexed where public access outweighs privacy ( [Google Spain](https://curia.europa.eu/jcms/upload/docs/application/pdf/2014-05/cp140070en.pdf); [GC and Others v. CNIL](https://curia.europa.eu/jcms/upload/docs/application/pdf/2019-09/cp190113en.pdf) ). ### How does NOSIBLE handle transparency when data comes from third-party sources? NOSIBLE collects from published sources rather than directly from every person named in those sources. Individual notice at web scale can be impossible or disproportionate, but [GDPR Article 14](https://gdpr-info.eu/art-14-gdpr/) still requires public information and safeguards. The [Article 29 Working Party transparency guidelines](https://www.edpb.europa.eu/system/files/2023-09/wp260rev01_en.pdf) require a documented analysis and public notice where Article 14(5)(b) is used. ### What happens to delivery files and customer outputs? Fast Search returns results in the HTTP response. Bulk Search and Time Search results are delivered as encrypted JSON files. Delivery setup controls customer-side access, retention, and handling. ### What privacy issues should customers consider with delivery files? Customers should review where result files land, who can access that storage, and how long downstream systems retain the outputs. NOSIBLE protects Bulk Search and Time Search delivery files with encryption, but customers control their downstream storage and access policies. Other legal and compliance policies [Compliance index](https://nosible.com/legal/compliance) [Crawling](https://nosible.com/legal/crawling) [Copyright](https://nosible.com/legal/copyright) [Security](https://nosible.com/legal/security) [Start Trial](https://nosible.com/start-trial) [Return to Trial Checklist](https://nosible.com/start-trial#legal) > How NOSIBLE minimizes personal data, retains limited usage logs, and handles removal, objection, erasure, and delisting requests. **URL:** https://nosible.com/legal/privacy --- --- title: "Our posture on security" description: "How NOSIBLE secures its APIs, infrastructure, data handling, delivery, access controls, and compliance work." url: "https://nosible.com/legal/security" --- Security # Our posture on security NOSIBLE security is built around dedicated infrastructure, hardened servers, controlled access, replication, and operational controls. The search index needs to stay available, consistent, recoverable, and protected from unwanted access under risk-based security standards such as [GDPR Article 32](https://gdpr-info.eu/art-32-gdpr/). NOSIBLE is operated by Nosible Inc. - NOSIBLE runs on dedicated [OVHcloud](https://www.ovhcloud.com/) bare-metal servers in France, Germany, and the UK. - Servers run hardened [Ubuntu Pro](https://ubuntu.com/pro) with controlled access and recorded activity. - Indexes are backed up to S3 and geographically replicated, with recovery paths kept separate. - Replication uses write-ahead logs, replica replay, and checksum validation. - Public API requests pass through a secure gateway and require API-key authentication, authorization, rate-limit checks, and validated inputs. - API traffic is encrypted in transit. Bulk Search and Time Search results are delivered as encrypted JSON files. - NOSIBLE does not log API request or response payloads on normal endpoints; limited usage logs are retained for 24 hours. - Server access is controlled through [Tailscale](https://tailscale.com/), MFA, and separated admin permissions. - NOSIBLE uses firewalls, endpoint security, antivirus, TLS, and DDoS protection. - Website verification checks security.txt, Google Web Risk, blocklists, and redirects. - NOSIBLE is in the early stages of SOC 2 readiness with Vanta and Cognisys. No independent examination or report has been completed. ## Security starts with dedicated infrastructure and controlled access NOSIBLE runs on dedicated bare-metal servers hosted by [OVHcloud](https://www.ovhcloud.com/) in France, Germany, and the UK. They run hardened [Ubuntu Pro](https://ubuntu.com/pro), sit behind firewalls and DDoS protection, and are administratively accessible only through [Tailscale](https://tailscale.com/) with MFA. Access and activity are recorded and auditable. The system is split into web discovery, web crawling, and web indexing. Discovery finds websites and URLs. Crawling retrieves HTML from whitelisted websites through Bright Data proxies and converts it into structured data. Raw HTML and parsed JSON are stored in Wasabi S3. Indexing streams structured data into lexical and vector indexes on OVH servers, using dedicated RunPod or Hugging Face GPU endpoints for embeddings. NOSIBLE World is a structured database distilled from the index. ## Index resilience is handled through backups and replication | Area | NOSIBLE position | | --- | --- | | Primary hosting | Dedicated [OVHcloud](https://www.ovhcloud.com/) bare-metal servers in France, Germany, and the UK. | | Backups | Indexes are backed up to S3 regularly. | | Replication | The index is geographically replicated using write-ahead logs shipped from primary indexes to replica indexes. | | Consistency checks | Replica indexes replay committed changes and use checksums to verify consistency after updates. | | Server hardening | Servers run [Ubuntu Pro](https://ubuntu.com/pro) with full hardening enabled. | | Access control | Server access is limited through [Tailscale](https://tailscale.com/), with access and activity recorded. | | Security program | NOSIBLE is in the early stages of SOC 2 readiness with Vanta and Cognisys. Policies are in place, vendors are mapped, and controls are being implemented. No independent examination or report has been completed. | ## Operational controls reduce intrusion and abuse risk Public API requests are proxied into the index through a secure gateway. Every request is authenticated with an API key and authorized against that key's permissions and rate limits. Inputs are validated against Pydantic models, rate-limit breaches return HTTP 429 responses, and retrieval endpoints are monitored for abuse. Behind the gateway, NOSIBLE uses firewalls, endpoint security, antivirus, anti-malware, multi-factor authentication, high-security Gmail settings, TLS encryption, DDoS protection, and hardened Ubuntu Pro. Administrative access requires Tailscale with MFA and is recorded and auditable. Secrets and per-key encryption keys are stored in [Google Secret Manager](https://cloud.google.com/security/products/secret-manager). U.S. security enforcement has focused on gaps between promised controls and actual controls, including [FTC v. Wyndham](https://www.ftc.gov/system/files/documents/cases/150824wyndhamopinion.pdf), [LabMD v. FTC](https://law.justia.com/cases/federal/appellate-courts/ca11/16-16270/16-16270-2018-06-06.html), [FTC Chegg](https://www.ftc.gov/news-events/news/press-releases/2022/10/ftc-brings-action-against-ed-tech-provider-chegg-careless-security-exposed-personal-data-millions), and [FTC Zoom](https://www.ftc.gov/news-events/news/press-releases/2020/11/ftc-requires-zoom-enhance-its-security-practices-part-settlement). ## Common security questions ### Where are NOSIBLE servers hosted? NOSIBLE runs on dedicated OVHcloud bare-metal servers in France, Germany, and the UK rather than on multi-tenant cloud compute. NOSIBLE also buys load balancers and networking products from OVHcloud. The core index benefits from predictable compute, storage, network behavior, and clearer operational control. ### How is NOSIBLE architected? NOSIBLE has three major systems: web discovery, web crawling, and web indexing. Discovery finds websites and URLs. Crawling retrieves HTML from whitelisted websites through Bright Data proxies and converts it into structured data. Raw HTML and parsed JSON are stored in Wasabi S3. Indexing streams structured data into lexical and vector indexes on dedicated OVHcloud servers, with embeddings supplied by dedicated RunPod or Hugging Face GPU endpoints. NOSIBLE World is a structured database distilled from the index. ### What deployment options does NOSIBLE support? NOSIBLE supports three deployment options: structured flat files delivered through S3 with no API or executable code; an on-premises Docker container that connects to S3 and exposes an internal search API; and managed Web API access. ### How are NOSIBLE APIs authenticated and authorized? NOSIBLE APIs use API keys. OAuth 2.0 and mTLS are not currently supported. Requests pass through a secure gateway and each key is checked for authentication, permissions, and rate limits. API keys can be revoked and reissued on request. ### How are API inputs and rate limits enforced? Inputs and payloads are validated against Pydantic models. When a rate limit is reached, the API returns an HTTP 429 rate-limit-exceeded response. Retrieval endpoints are strictly rate-limited and monitored for abuse. ### Is data encrypted in transit and at rest? All requests to and from NOSIBLE APIs are encrypted in transit with TLS 1.2 or newer. Bulk Search and Time Search results are delivered as encrypted JSON files. The live search index is not encrypted at rest because it must remain queryable; NOSIBLE mitigates that exposure with dedicated infrastructure, firewalls, DDoS protection, restricted administrative access through Tailscale with MFA, and recorded access activity. ### What customer API data does NOSIBLE retain? NOSIBLE retains limited usage logs for 24 hours to enforce rate limits and prevent abuse. It does not log or record API request or response payloads on normal endpoints. Customer API requests are processed on OVHcloud servers in France, Germany, and the UK and are not stored. ### Where are delivery files and supporting services hosted? Raw crawl artifacts are stored in Wasabi S3. Google Cloud supports datastore, storage, error reporting, and secret management. NOSIBLE World S3 delivery is served from AWS us-east-1. Bulk Search and Time Search results are delivered as encrypted JSON files. ### What logging and monitoring does NOSIBLE use? Usage logs retain only what is needed for rate limits and abuse prevention and are deleted after 24 hours. API payloads and customer content are not included in logs. Administrative server access and activity through Tailscale are recorded and auditable. NOSIBLE is in the early stages of using Vanta for vendor and control monitoring. ### Have NOSIBLE APIs undergone security testing? Internal code reviews have occurred, but NOSIBLE has not yet completed penetration testing or an independent vulnerability assessment. Vulnerability and control monitoring through Vanta is still in the early stages. NOSIBLE will share an executive summary after independent testing is completed. ### How does NOSIBLE keep replicas consistent? The index uses write-ahead log replication. Committed changes are shipped from a primary index to replica indexes. Replica indexes replay those logs to stay synchronized, and checksums verify that the index is consistent across replicas after updates. Replication protects availability and integrity under [GDPR Article 32](https://gdpr-info.eu/art-32-gdpr/). ### What does NOSIBLE back up? NOSIBLE backs up indexes to S3 regularly and keeps the index geographically replicated. Raw HTML and parsed JSON documents are stored in secure S3 buckets hosted by Wasabi. These backups and object stores support recovery, durability, replay, vendor review, and continuity if a primary system fails. ### What operating system hardening is used? NOSIBLE servers use [Ubuntu Pro](https://ubuntu.com/pro) with full hardening enabled. NOSIBLE also uses server firewalls, endpoint security, antivirus, anti-malware, multi-factor authentication, email security controls, SSL encryption, and DDoS protection for the primary index and replicas. These controls reduce intrusion risk across servers, endpoints, and network access. ### How does website verification reduce security risk? Before a website enters the source universe, NOSIBLE checks whether the homepage resolves cleanly, whether it redirects to a different entity, whether it appears on internal blocklists, whether Google Web Risk reports malware, social engineering, or unwanted software, and whether the site publishes security.txt. These checks reduce unsafe-source risk before content reaches the index. ### Why does NOSIBLE check security.txt? security.txt can identify a website's vulnerability disclosure channel. NOSIBLE looks for it at the standard root and well-known paths and records the path when present. That does not prove a site is secure, but it gives a useful operational contact signal when a source needs review or when a security issue needs to be routed responsibly. ### How is server access controlled? Servers are accessible only through [Tailscale](https://tailscale.com/), and access and activity are recorded. Access to core NOSIBLE systems also uses multi-factor authentication. Administrative access is separate from customer API permissions and governs the infrastructure that runs the index, replicas, storage, and supporting services behind customer-facing products. ### Does NOSIBLE claim to be SOC 2 certified? No. SOC 2 is an independent attestation report, not a certification. NOSIBLE is in the early stages of readiness work with Vanta and Cognisys. Policies are in place, vendors have been mapped, and controls are being implemented, but no independent examination or report has been completed. ### How would NOSIBLE handle breach notification? NOSIBLE would assess notification duties under the laws that apply to the affected customers and data. UK and EU GDPR require supervisory-authority notification without undue delay and, where feasible, within 72 hours after awareness of a qualifying personal-data breach ( [GDPR Article 33](https://gdpr-info.eu/art-33-gdpr/) ). Data-subject notice applies where the breach is likely to create high risk ( [GDPR Article 34](https://gdpr-info.eu/art-34-gdpr/) ). U.S. timelines vary by state. ### Which third-party vendors does NOSIBLE rely on? NOSIBLE uses [Bright Data](https://brightdata.com/) for proxy infrastructure, [OVHcloud](https://www.ovhcloud.com/) for bare-metal hosting, [Hugging Face](https://huggingface.co/) and [RunPod](https://www.runpod.io/) for dedicated GPU-hosted embedding endpoints, [Google Cloud](https://cloud.google.com/) for datastore, storage, error reporting, and secret management, [Wasabi](https://wasabi.com/) for raw HTML and parsed JSON in S3-compatible storage, and AWS us-east-1 for NOSIBLE World S3 delivery. NOSIBLE maintains alternatives for critical vendors so it can move without redesigning the service. ### What open-source technology does NOSIBLE use? NOSIBLE is built with Python and runs in Docker containers on bare-metal OVH servers. Core packages used across the system include NumPy, Polars, Numba, Redis bindings, orjson, Zstandard, Aho-Corasick tools, spaCy, wordfreq, PyStemmer, fast language detection, SimSIMD, Flask, Gunicorn, FastAPI, Uvicorn, BeautifulSoup, LMDB bindings, OpenAI tooling, rbloom, datefinder, sentence-transformers, json repair tools, SciPy, scikit-learn, usearch, and Playwright. ### Which embedding models and model vendors does NOSIBLE use? NOSIBLE Search uses Microsoft's multilingual-e5-large-instruct model through dedicated Hugging Face and RunPod GPU endpoints. NOSIBLE World uses OpenAI's text-embedding-3-large model. ### Is NOSIBLE SOC 2 compliant? Not yet. NOSIBLE is in the early stages of SOC 2 readiness with Vanta and Cognisys. Policies are in place, vendors have been mapped, and controls are being implemented. No independent examination has been completed and there is no SOC 2 report to share today. In the meantime, NOSIBLE can share its security posture, policies, and compliance pack during customer due diligence. Other legal and compliance policies [Compliance index](https://nosible.com/legal/compliance) [Crawling](https://nosible.com/legal/crawling) [Copyright](https://nosible.com/legal/copyright) [Privacy](https://nosible.com/legal/privacy) [Start Trial](https://nosible.com/start-trial) [Return to Trial Checklist](https://nosible.com/start-trial#legal) > How NOSIBLE secures its APIs, infrastructure, data handling, delivery, access controls, and compliance work. **URL:** https://nosible.com/legal/security --- --- title: "NOSIBLE Vendor Comparisons" description: "Compare NOSIBLE SEARCH and WORLD with financial news, ESG risk, media intelligence, and open event data vendors across coverage, timing, APIs, and model workflows." url: "https://nosible.com/compare" --- Compare / Updated July 2026 # NOSIBLE Vendor Comparisons Compare NOSIBLE WORLD with the financial news, ESG risk, media intelligence, and open event datasets that buyers evaluate most often. [Market intelligence RavenPack Financial-news analytics, sentiment signals, source evidence, and historical research workflows.](https://nosible.com/compare/nosible-vs-ravenpack) [Market intelligence Bloomberg Market news feeds, open-web context, research agents, and historical evidence.](https://nosible.com/compare/nosible-vs-bloomberg) [Market intelligence LSEG Machine-readable news, market events, source evidence, and historical research workflows.](https://nosible.com/compare/nosible-vs-lseg) [Market intelligence Dow Jones Enterprise news intelligence, source access, language coverage, and research workflows.](https://nosible.com/compare/nosible-vs-dow-jones) [Market intelligence AlphaSense Market research, licensed sources, AI workflows, and historical analysis.](https://nosible.com/compare/nosible-vs-alphasense) [Market intelligence Kensho Financial AI data, S&P access, cited retrieval, and ranked events.](https://nosible.com/compare/nosible-vs-kensho) [ESG & risk FactSet Outside-in ESG signals, issuer context, source evidence, and research workflows.](https://nosible.com/compare/nosible-vs-factset-truvalue) [ESG & risk RepRisk ESG risk data, business-conduct signals, source evidence, and research workflows.](https://nosible.com/compare/nosible-vs-reprisk) [ESG & risk SESAMm ESG intelligence, corporate signals, point-in-time evidence, and agent workflows.](https://nosible.com/compare/nosible-vs-sesamm) [ESG & risk Signal AI Reputation intelligence, risk signals, source coverage, and research workflows.](https://nosible.com/compare/nosible-vs-signal-ai) [Media intelligence MarketPsych Sentiment analytics, behavioral signals, source evidence, and agent workflows.](https://nosible.com/compare/nosible-vs-marketpsych) [Media intelligence MKT MediaStats Point-in-time media data, raw sources, factor workflows, and backtesting.](https://nosible.com/compare/nosible-vs-mkt-mediastats) [Media intelligence Opoint Media intelligence, curated news data, source evidence, and research workflows.](https://nosible.com/compare/nosible-vs-opoint) [Media intelligence Alexandria Technology AI-ready research, point-in-time retrieval, sentiment data, and source coverage.](https://nosible.com/compare/nosible-vs-alexandria) [Agent & web research Exa Web-search APIs, agent research, source evidence, and structured web data.](https://nosible.com/compare/nosible-vs-exa) [Agent & web research Parallel Web Systems Web-research APIs, task execution, cited sources, and agent workflows.](https://nosible.com/compare/nosible-vs-parallel) [Agent & web research Tavily Agent search, extraction, crawling, source evidence, and research delivery.](https://nosible.com/compare/nosible-vs-tavily) [Agent & web research News API News search, headline feeds, source coverage, and research workflows.](https://nosible.com/compare/nosible-vs-news-api) [Open data GDELT Global-news data, event history, open sources, and research workflows.](https://nosible.com/compare/nosible-vs-gdelt) [Open data Common Crawl Open-web archives, crawl indexes, source evidence, and processing tradeoffs.](https://nosible.com/compare/nosible-vs-common-crawl) > Compare NOSIBLE SEARCH and WORLD with financial news, ESG risk, media intelligence, and open event data vendors across coverage, timing, APIs, and model workflows. **URL:** https://nosible.com/compare --- --- title: "NOSIBLE vs RavenPack" description: "Compare NOSIBLE and RavenPack on point-in-time event intelligence, source coverage, sentiment, delivery, AI agents, and backtesting." url: "https://nosible.com/compare/nosible-vs-ravenpack" --- Comparison / Reviewed July 14, 2026 # NOSIBLE vs RavenPack RavenPack structures financial news into sentiment, relevance, novelty, and event data. [[ 2 a]](https://archive.is/pnQmK) [[ 2 b]](https://web.archive.org/web/20260509203524/https://www.ravenpack.com/blog/machine-readable-news) NOSIBLE is built around dated open-web retrieval, ranked events, and agent workflows. [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) NOSIBLE-AUTHORED COMPARISON · NOSIBLE HAS A COMMERCIAL INTEREST IN THIS COMPARISON · FACTS ATTRIBUTED TO FIRST-PARTY VENDOR MATERIALS · EVALUATIVE STATEMENTS ARE NOSIBLE'S OPINION · REVIEWED JULY 14, 2026 · DUAL PUBLIC ARCHIVES WHERE SUPPORTED · [STANDARDS & CORRECTIONS](https://nosible.com/compare/nosible-vs-ravenpack#comparison-standards). IF YOU REPRESENT RAVENPACK AND BELIEVE A FACTUAL STATEMENT IS INACCURATE, EMAIL [STUART@NOSIBLE.COM](mailto:stuart@nosible.com?subject=Correction%20request%3A%20NOSIBLE%20vs%20RavenPack) WITH THE SPECIFIC CLAIM AND A SUPPORTING FIRST-PARTY URL. NOSIBLE WILL REVIEW AND CORRECT SUBSTANTIATED ERRORS. - RavenPack's cited Edge News Analytics page says it processes more than 40,000 news and social-media sources in 13 languages. [[ 1 a]](https://archive.is/s95Ec) [[ 1 b]](https://web.archive.org/web/20260311110344/https://www.ravenpack.com/products/edge/data/news-analytics) - RavenPack reports 12 million named entities and 7,000+ event topics; NOSIBLE WORLD v1.2 reports about 3.2 million organizations and 6.4 million people. [[ 5 a]](https://archive.is/tGYc9) [[ 6 a]](https://archive.is/tnkpG) Vendor definitions may differ. [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) - Its published analytics include sentiment, relevance, novelty, temporal, and impact measures. [[ 2 a]](https://archive.is/pnQmK) [[ 2 b]](https://web.archive.org/web/20260509203524/https://www.ravenpack.com/blog/machine-readable-news) - NOSIBLE emphasizes dated source retrieval, ranked events, multilingual open-web coverage, and agent access. [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) - The products' published positioning differs: RavenPack offers packaged finance-native analytics, while NOSIBLE offers a source-and-event layer. [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) - In NOSIBLE's view, consider NOSIBLE when inspectable source evidence and historical event retrieval are central to the workflow. ## Different published starting points RavenPack starts with large-scale financial-news analytics: it identifies entities and events, then supplies measures such as sentiment, relevance, novelty, and impact. [[ 2 a]](https://archive.is/pnQmK) [[ 2 b]](https://web.archive.org/web/20260509203524/https://www.ravenpack.com/blog/machine-readable-news) NOSIBLE starts with dated source retrieval and ranked open-web events for agents, research systems, and historical analysis. [[ 4 a]](https://archive.is/glqrw) In NOSIBLE's view, NOSIBLE is the stronger fit when a team wants to inspect and reuse the underlying evidence rather than begin with a finished analytics feed. ## News analytics and source-and-event workflows The competitor column summarizes what the named vendor's cited first-party materials describe; the NOSIBLE column summarizes NOSIBLE's cited materials. Primary use NOSIBLE Point-in-time open-web intelligence [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) RavenPack Systematic news analytics for financial workflows [[ 1 a]](https://archive.is/s95Ec) [[ 1 b]](https://web.archive.org/web/20260311110344/https://www.ravenpack.com/products/edge/data/news-analytics) Illustrative users NOSIBLE AI agents, backtests, risk systems [[ 4 a]](https://archive.is/glqrw) RavenPack Quantitative and discretionary investment teams [[ 1 a]](https://archive.is/s95Ec) [[ 1 b]](https://web.archive.org/web/20260311110344/https://www.ravenpack.com/products/edge/data/news-analytics) Sources NOSIBLE 300,000+ open-web sources [[ 4 a]](https://archive.is/glqrw) RavenPack Cited Edge page: 40,000+ news and social-media sources [[ 1 a]](https://archive.is/s95Ec) [[ 1 b]](https://web.archive.org/web/20260311110344/https://www.ravenpack.com/products/edge/data/news-analytics) Languages NOSIBLE 95 [[ 4 a]](https://archive.is/glqrw) RavenPack Cited Edge page: 13 languages [[ 1 a]](https://archive.is/s95Ec) [[ 1 b]](https://web.archive.org/web/20260311110344/https://www.ravenpack.com/products/edge/data/news-analytics) History NOSIBLE About 30 years [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) RavenPack RavenPack says it has developed financial sentiment analysis since 2003 [[ 1 a]](https://archive.is/s95Ec) [[ 1 b]](https://web.archive.org/web/20260311110344/https://www.ravenpack.com/products/edge/data/news-analytics) Point-in-time method NOSIBLE Five-way date verification [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) RavenPack Temporal scoring and historical analytics [[ 1 a]](https://archive.is/s95Ec) [[ 1 b]](https://web.archive.org/web/20260311110344/https://www.ravenpack.com/products/edge/data/news-analytics) Event representation NOSIBLE 100M+ ranked dated events [[ 5 a]](https://archive.is/tGYc9) RavenPack Cited Edge page: 7,000+ event topics [[ 1 a]](https://archive.is/s95Ec) [[ 1 b]](https://web.archive.org/web/20260311110344/https://www.ravenpack.com/products/edge/data/news-analytics) Source model NOSIBLE Open-web long-form sources; excludes social, paywalled, marketplaces, and adult content [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) RavenPack Premium news, web content, and social media [[ 2 a]](https://archive.is/pnQmK) [[ 2 b]](https://web.archive.org/web/20260509203524/https://www.ravenpack.com/blog/machine-readable-news) Delivery NOSIBLE Search API, World, agents, SDKs, MCP [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) RavenPack Structured document- and event-level analytics [[ 1 a]](https://archive.is/s95Ec) [[ 1 b]](https://web.archive.org/web/20260311110344/https://www.ravenpack.com/products/edge/data/news-analytics) Pricing NOSIBLE Free tier, higher tiers by request [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) RavenPack Trial and commercial access through RavenPack [[ 1 a]](https://archive.is/s95Ec) [[ 1 b]](https://web.archive.org/web/20260311110344/https://www.ravenpack.com/products/edge/data/news-analytics) | Dimension | NOSIBLE | RavenPack | | --- | --- | --- | | Primary use | Point-in-time open-web intelligence [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Systematic news analytics for financial workflows [[ 1 a]](https://archive.is/s95Ec) [[ 1 b]](https://web.archive.org/web/20260311110344/https://www.ravenpack.com/products/edge/data/news-analytics) | | Illustrative users | AI agents, backtests, risk systems [[ 4 a]](https://archive.is/glqrw) | Quantitative and discretionary investment teams [[ 1 a]](https://archive.is/s95Ec) [[ 1 b]](https://web.archive.org/web/20260311110344/https://www.ravenpack.com/products/edge/data/news-analytics) | | Sources | 300,000+ open-web sources [[ 4 a]](https://archive.is/glqrw) | Cited Edge page: 40,000+ news and social-media sources [[ 1 a]](https://archive.is/s95Ec) [[ 1 b]](https://web.archive.org/web/20260311110344/https://www.ravenpack.com/products/edge/data/news-analytics) | | Languages | 95 [[ 4 a]](https://archive.is/glqrw) | Cited Edge page: 13 languages [[ 1 a]](https://archive.is/s95Ec) [[ 1 b]](https://web.archive.org/web/20260311110344/https://www.ravenpack.com/products/edge/data/news-analytics) | | History | About 30 years [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | RavenPack says it has developed financial sentiment analysis since 2003 [[ 1 a]](https://archive.is/s95Ec) [[ 1 b]](https://web.archive.org/web/20260311110344/https://www.ravenpack.com/products/edge/data/news-analytics) | | Point-in-time method | Five-way date verification [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Temporal scoring and historical analytics [[ 1 a]](https://archive.is/s95Ec) [[ 1 b]](https://web.archive.org/web/20260311110344/https://www.ravenpack.com/products/edge/data/news-analytics) | | Event representation | 100M+ ranked dated events [[ 5 a]](https://archive.is/tGYc9) | Cited Edge page: 7,000+ event topics [[ 1 a]](https://archive.is/s95Ec) [[ 1 b]](https://web.archive.org/web/20260311110344/https://www.ravenpack.com/products/edge/data/news-analytics) | | Source model | Open-web long-form sources; excludes social, paywalled, marketplaces, and adult content [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Premium news, web content, and social media [[ 2 a]](https://archive.is/pnQmK) [[ 2 b]](https://web.archive.org/web/20260509203524/https://www.ravenpack.com/blog/machine-readable-news) | | Delivery | Search API, World, agents, SDKs, MCP [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) | Structured document- and event-level analytics [[ 1 a]](https://archive.is/s95Ec) [[ 1 b]](https://web.archive.org/web/20260311110344/https://www.ravenpack.com/products/edge/data/news-analytics) | | Pricing | Free tier, higher tiers by request [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Trial and commercial access through RavenPack [[ 1 a]](https://archive.is/s95Ec) [[ 1 b]](https://web.archive.org/web/20260311110344/https://www.ravenpack.com/products/edge/data/news-analytics) | ## Potential fit by workflow RavenPack publishes ready-made analytics over a finance-oriented news universe. [[ 1 a]](https://archive.is/s95Ec) [[ 1 b]](https://web.archive.org/web/20260311110344/https://www.ravenpack.com/products/edge/data/news-analytics) NOSIBLE is designed for the adjacent problem: retrieving dated source material across the open web, ranking the associated events, and making that evidence available to agents and historical research. [[ 4 a]](https://archive.is/glqrw) In NOSIBLE's view, the two approaches may be used separately or together. ## Common RavenPack comparison questions ### How does NOSIBLE feel it differentiates itself from RavenPack? NOSIBLE is an AI-native company with two products: SEARCH and WORLD. [[ 5 a]](https://archive.is/tGYc9) [[ 7 a]](https://archive.is/iDZeo) SEARCH lets agents find dated open-web sources they can cite and inspect directly. [[ 4 a]](https://archive.is/glqrw) WORLD is a live open-web event database for models and backtests, with an embedding per event. [[ 5 a]](https://archive.is/tGYc9) [[ 6 a]](https://archive.is/tnkpG) [[ 8 a]](https://archive.is/1CUvQ) NOSIBLE is committed to open-source software and makes its models publicly available on Hugging Face. [[ 9 a]](https://archive.is/7SSz0) [[ 10 a]](https://archive.is/kHxMG) Related [WORLD v1.2 trial](https://nosible.com/start-trial#data-coverage) [Sentiment model](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) [Forward-looking model](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) ### If we already use RavenPack Edge or Bigdata.com, what gap would NOSIBLE fill? NOSIBLE adds a source-and-event layer around packaged news analytics. [[ 1 a]](https://archive.is/s95Ec) [[ 1 b]](https://web.archive.org/web/20260311110344/https://www.ravenpack.com/products/edge/data/news-analytics) [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) It is intended for agents and researchers that need dated documents, ranked events, entity context, and source inspection across the open web. [[ 1 a]](https://archive.is/s95Ec) [[ 1 b]](https://web.archive.org/web/20260311110344/https://www.ravenpack.com/products/edge/data/news-analytics) [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) RavenPack publishes sentiment and event analytics. [[ 1 a]](https://archive.is/s95Ec) [[ 1 b]](https://web.archive.org/web/20260311110344/https://www.ravenpack.com/products/edge/data/news-analytics) In NOSIBLE's view, RavenPack remains the more direct fit when those finished analytics are the required output. Related [Sentiment model](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) [Sentiment research](https://nosible.com/blog/fast-enough-to-matter-productionizing-tiny-transformers-for-signal-extraction) [WORLD event database](https://nosible.world/world) ### How does NOSIBLE compare with RavenPack's premium-source and external-search model? RavenPack describes a mix of premium publishers, web content, and social media. [[ 2 a]](https://archive.is/pnQmK) [[ 2 b]](https://web.archive.org/web/20260509203524/https://www.ravenpack.com/blog/machine-readable-news) NOSIBLE uses an open-web retrieval model and emphasizes replayable source history. [[ 4 a]](https://archive.is/glqrw) Buyers should test both products against their own source list, languages, entitlements, and historical-query requirements. Related [WORLD event database](https://nosible.world/world) ### Do we lose RavenPack's entity mapping and event taxonomy if we switch? RavenPack's cited Edge page publishes a taxonomy of more than 7,000 event topics and coverage of more than 12 million named entities. [[ 1 a]](https://archive.is/s95Ec) [[ 1 b]](https://web.archive.org/web/20260311110344/https://www.ravenpack.com/products/edge/data/news-analytics) NOSIBLE supplies its own event, entity, ticker, and risk mappings, but it is not a drop-in copy of RavenPack's schema. [[ 1 a]](https://archive.is/s95Ec) [[ 1 b]](https://web.archive.org/web/20260311110344/https://www.ravenpack.com/products/edge/data/news-analytics) [[ 5 a]](https://archive.is/tGYc9) A migration should therefore include a field-level mapping exercise. Related [WORLD event database](https://nosible.world/world) ### What if our workflow depends on licensed filings, transcripts, or earnings content? NOSIBLE's cited materials describe open-web retrieval, not a license to a specialist transcript feed. [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) They describe dated government, company, regional, and specialist web sources connected to events and entities. [[ 1 a]](https://archive.is/s95Ec) [[ 1 b]](https://web.archive.org/web/20260311110344/https://www.ravenpack.com/products/edge/data/news-analytics) [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) In NOSIBLE's view, NOSIBLE may provide contextual evidence around separately licensed content. Teams should retain specialist content where its rights and coverage are required. Related [Copyright posture](https://nosible.com/legal/copyright) [WORLD event database](https://nosible.world/world) ### How should a quant evaluate point-in-time claims from both vendors? Ask each vendor what its timestamps represent, how revisions are handled, and what a past-date query can return. Then run a fixed historical test set. [[ 1 a]](https://archive.is/s95Ec) [[ 1 b]](https://web.archive.org/web/20260311110344/https://www.ravenpack.com/products/edge/data/news-analytics) [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) NOSIBLE is designed around replayable dated sources and events; RavenPack publishes temporal and historical analytics that should be evaluated against the same protocol. [[ 1 a]](https://archive.is/s95Ec) [[ 1 b]](https://web.archive.org/web/20260311110344/https://www.ravenpack.com/products/edge/data/news-analytics) [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Related [WORLD event database](https://nosible.world/world) Continue comparing ## Related Vendor Comparisons Compare RavenPack with adjacent options for task execution, agent tooling, data access, and research workflows, then evaluate the output shape your application actually needs for production. [Bloomberg](https://nosible.com/compare/nosible-vs-bloomberg) [LSEG](https://nosible.com/compare/nosible-vs-lseg) [Dow Jones](https://nosible.com/compare/nosible-vs-dow-jones) [GDELT](https://nosible.com/compare/nosible-vs-gdelt) Dive deeper ## Take your next step today Review the delivered field definitions, classification boundaries, and example values before comparing vendor workflows, so your team can assess what each product actually returns to downstream systems. [Data dictionaries](https://nosible.com/data-dictionaries) [Ontology reference](https://nosible.com/ontologies) [API reference](https://docs.nosible.com/) [Start trial](https://nosible.com/start-trial) Sources reviewed July 14, 2026 : [[ 1 ] RavenPack News Analytics](https://www.ravenpack.com/products/edge/data/news-analytics) ( [dated snapshot](https://archive.is/s95Ec); [Wayback copy](https://web.archive.org/web/20260311110344/https://www.ravenpack.com/products/edge/data/news-analytics) ) , [[ 2 ] RavenPack machine-readable news overview](https://www.ravenpack.com/blog/machine-readable-news) ( [dated snapshot](https://archive.is/pnQmK); [Wayback copy](https://web.archive.org/web/20260509203524/https://www.ravenpack.com/blog/machine-readable-news) ) , [[ 3 ] NOSIBLE product overview](https://nosible.com/) ( [dated snapshot](https://archive.is/r7qDv); [Wayback copy](https://web.archive.org/web/20260713185114/https://nosible.com/) ) , [[ 4 ] NOSIBLE SEARCH](https://docs.nosible.com/) ( [dated snapshot](https://archive.is/glqrw) ) , [[ 5 ] NOSIBLE WORLD](https://nosible.world/world) ( [dated snapshot](https://archive.is/tGYc9) ) , [[ 6 ] NOSIBLE WORLD v1.2 trial and coverage](https://nosible.com/start-trial) ( [dated snapshot](https://archive.is/tnkpG) ) , [[ 7 ] NOSIBLE AI-native research overview](https://nosible.com/blog) ( [dated snapshot](https://archive.is/iDZeo) ) , [[ 8 ] NOSIBLE embedding-based research](https://nosible.com/blog/an-embedding-based-approach-to-trade-and-economic-policy-uncertainty) ( [dated snapshot](https://archive.is/1CUvQ) ) , [[ 9 ] NOSIBLE Financial Sentiment v1.2 Base](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) ( [dated snapshot](https://archive.is/7SSz0) ) , [[ 10 ] NOSIBLE Forward-Looking v1.2 Base](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) ( [dated snapshot](https://archive.is/kHxMG) ) . RavenPack, RavenPack Edge, and Bigdata.com are trademarks of their respective owners. NOSIBLE is not affiliated with, sponsored by, or endorsed by RavenPack. Comparison standards, legal context & corrections This comparison was prepared by NOSIBLE, which has a commercial interest in the products being compared. It is based on the cited public materials as they appeared on July 14, 2026. NOSIBLE has not tested every competitor feature, and the page is not a complete statement of either product. Products change; confirm current requirements, availability, and commercial terms with each vendor. Factual claims are attributed to cited first-party materials; evaluative statements reflect NOSIBLE's opinion. Each source includes a dated archive.is snapshot and, where the Internet Archive captured that URL, a timestamp-specific Wayback copy. Archive availability is controlled by those services. Competitor names and marks are used only to identify the products being compared. No affiliation, sponsorship, or endorsement is implied. The FTC says truthful, non-deceptive comparative advertising may identify competitors ( [policy](https://www.ftc.gov/legal-library/browse/statement-policy-regarding-comparative-advertising); [dated snapshot](https://archive.is/nZXY4); [Wayback copy](https://web.archive.org/web/20260215042003/https://www.ftc.gov/legal-library/browse/statement-policy-regarding-comparative-advertising) ). For one U.S. example of nominative-use analysis, see *New Kids on the Block v. News America Publishing, Inc.*, 971 F.2d 302, 308 (9th Cir. 1992) ( [opinion](https://law.justia.com/cases/federal/appellate-courts/F2/971/302/72076/); [dated snapshot](https://archive.is/dnk8j); [Wayback copy](https://web.archive.org/web/20250211213616/https://law.justia.com/cases/federal/appellate-courts/F2/971/302/72076/) ). If you represent RavenPack and believe a factual statement is inaccurate, email [stuart@nosible.com](mailto:stuart@nosible.com?subject=Correction%20request%3A%20NOSIBLE%20vs%20RavenPack) with the specific claim and a supporting first-party URL. NOSIBLE will review and correct substantiated errors. > Compare NOSIBLE and RavenPack on point-in-time event intelligence, source coverage, sentiment, delivery, AI agents, and backtesting. **URL:** https://nosible.com/compare/nosible-vs-ravenpack --- --- title: "NOSIBLE vs Bloomberg" description: "Compare NOSIBLE and Bloomberg for point-in-time event intelligence, open-web search, AI agents, backtesting, multilingual coverage, delivery, and market signal discovery." url: "https://nosible.com/compare/nosible-vs-bloomberg" --- Comparison / Reviewed July 14, 2026 # NOSIBLE vs Bloomberg Bloomberg offers real-time, machine-readable news and market-data feeds. [[ 1 a]](https://archive.is/ucKET) [[ 1 b]](https://web.archive.org/web/20260512034507/https://professional.bloomberg.com/products/data/enterprise-catalog/event-driven-feeds/) NOSIBLE focuses on dated open-web source retrieval, ranked events, and AI-agent workflows. [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) NOSIBLE-AUTHORED COMPARISON · NOSIBLE HAS A COMMERCIAL INTEREST IN THIS COMPARISON · FACTS ATTRIBUTED TO FIRST-PARTY VENDOR MATERIALS · EVALUATIVE STATEMENTS ARE NOSIBLE'S OPINION · REVIEWED JULY 14, 2026 · DUAL PUBLIC ARCHIVES WHERE SUPPORTED · [STANDARDS & CORRECTIONS](https://nosible.com/compare/nosible-vs-bloomberg#comparison-standards). IF YOU REPRESENT BLOOMBERG AND BELIEVE A FACTUAL STATEMENT IS INACCURATE, EMAIL [STUART@NOSIBLE.COM](mailto:stuart@nosible.com?subject=Correction%20request%3A%20NOSIBLE%20vs%20Bloomberg) WITH THE SPECIFIC CLAIM AND A SUPPORTING FIRST-PARTY URL. NOSIBLE WILL REVIEW AND CORRECT SUBSTANTIATED ERRORS. - Bloomberg describes Event-Driven Feeds as structured, real-time, machine-readable data for black-box applications. [[ 1 a]](https://archive.is/ucKET) [[ 1 b]](https://web.archive.org/web/20260512034507/https://professional.bloomberg.com/products/data/enterprise-catalog/event-driven-feeds/) - Its textual-news feed combines Bloomberg reporting with selected third-party, web, and social sources. [[ 2 a]](https://archive.is/rarjM) [[ 2 b]](https://web.archive.org/web/20260713190534/https://assets.bbhub.io/professional/sites/41/Fact-Sheet-EDF-Textual-News.pdf) - Bloomberg News Analytics includes sentiment, novelty, readership-heat, and social-velocity measures. [[ 1 a]](https://archive.is/ucKET) [[ 1 b]](https://web.archive.org/web/20260512034507/https://professional.bloomberg.com/products/data/enterprise-catalog/event-driven-feeds/) - NOSIBLE provides dated source retrieval, ranked events, multilingual coverage, and agent-oriented access. [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) - In NOSIBLE's view, consider NOSIBLE when the workflow begins with inspectable open-web evidence rather than a market-data feed. ## Feed infrastructure and source evidence Bloomberg's published Event-Driven Feeds move structured news, analytics, economic indicators, and corporate-event data into systematic applications in real time. [[ 1 a]](https://archive.is/ucKET) [[ 1 b]](https://web.archive.org/web/20260512034507/https://professional.bloomberg.com/products/data/enterprise-catalog/event-driven-feeds/) NOSIBLE is designed for agents and researchers that need to search dated open-web material, inspect the source record, and retrieve ranked events. [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) In NOSIBLE's view, NOSIBLE is the stronger fit when the evidence itself is the product requirement. ## Structured feeds and dated source evidence The competitor column summarizes what the named vendor's cited first-party materials describe; the NOSIBLE column summarizes NOSIBLE's cited materials. Primary use NOSIBLE Point-in-time open-web event intelligence [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Bloomberg Real-time machine-readable news and financial data [[ 1 a]](https://archive.is/ucKET) [[ 1 b]](https://web.archive.org/web/20260512034507/https://professional.bloomberg.com/products/data/enterprise-catalog/event-driven-feeds/) Illustrative users NOSIBLE AI agents, backtests, research, risk systems [[ 4 a]](https://archive.is/glqrw) Bloomberg Systematic and enterprise market-data teams [[ 1 a]](https://archive.is/ucKET) [[ 1 b]](https://web.archive.org/web/20260512034507/https://professional.bloomberg.com/products/data/enterprise-catalog/event-driven-feeds/) Content scope NOSIBLE Long-form news, corporate, and government text [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Bloomberg Textual news, news analytics, economic data, and corporate-event feeds [[ 1 a]](https://archive.is/ucKET) [[ 1 b]](https://web.archive.org/web/20260512034507/https://professional.bloomberg.com/products/data/enterprise-catalog/event-driven-feeds/) Source types NOSIBLE Open-web and institutional text [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Bloomberg Bloomberg reporting plus selected third-party, web, and social content [[ 2 a]](https://archive.is/rarjM) [[ 2 b]](https://web.archive.org/web/20260713190534/https://assets.bbhub.io/professional/sites/41/Fact-Sheet-EDF-Textual-News.pdf) Published source metric NOSIBLE 300,000+ open-web sources [[ 4 a]](https://archive.is/glqrw) Bloomberg Cited 2015 fact sheet: more than 100,000 sources [[ 2 a]](https://archive.is/rarjM) [[ 2 b]](https://web.archive.org/web/20260713190534/https://assets.bbhub.io/professional/sites/41/Fact-Sheet-EDF-Textual-News.pdf) History NOSIBLE About 30 years [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Bloomberg Cited 2015 fact sheet: textual-news archive dating to 1992 [[ 2 a]](https://archive.is/rarjM) [[ 2 b]](https://web.archive.org/web/20260713190534/https://assets.bbhub.io/professional/sites/41/Fact-Sheet-EDF-Textual-News.pdf) Point-in-time method NOSIBLE Five-way date verification [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Bloomberg Real-time structured delivery [[ 1 a]](https://archive.is/ucKET) [[ 1 b]](https://web.archive.org/web/20260512034507/https://professional.bloomberg.com/products/data/enterprise-catalog/event-driven-feeds/) Events NOSIBLE 100M+ ranked dated events [[ 5 a]](https://archive.is/tGYc9) Bloomberg Corporate-event calendar and market-event data [[ 1 a]](https://archive.is/ucKET) [[ 1 b]](https://web.archive.org/web/20260512034507/https://professional.bloomberg.com/products/data/enterprise-catalog/event-driven-feeds/) Delivery NOSIBLE API, SDKs, MCP, Cybernaut-1 [[ 4 a]](https://archive.is/glqrw) Bloomberg Machine-readable enterprise feeds [[ 1 a]](https://archive.is/ucKET) [[ 1 b]](https://web.archive.org/web/20260512034507/https://professional.bloomberg.com/products/data/enterprise-catalog/event-driven-feeds/) Pricing NOSIBLE Product access through NOSIBLE [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Bloomberg Contact Bloomberg for access and commercial terms [[ 1 a]](https://archive.is/ucKET) [[ 1 b]](https://web.archive.org/web/20260512034507/https://professional.bloomberg.com/products/data/enterprise-catalog/event-driven-feeds/) | Dimension | NOSIBLE | Bloomberg | | --- | --- | --- | | Primary use | Point-in-time open-web event intelligence [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Real-time machine-readable news and financial data [[ 1 a]](https://archive.is/ucKET) [[ 1 b]](https://web.archive.org/web/20260512034507/https://professional.bloomberg.com/products/data/enterprise-catalog/event-driven-feeds/) | | Illustrative users | AI agents, backtests, research, risk systems [[ 4 a]](https://archive.is/glqrw) | Systematic and enterprise market-data teams [[ 1 a]](https://archive.is/ucKET) [[ 1 b]](https://web.archive.org/web/20260512034507/https://professional.bloomberg.com/products/data/enterprise-catalog/event-driven-feeds/) | | Content scope | Long-form news, corporate, and government text [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Textual news, news analytics, economic data, and corporate-event feeds [[ 1 a]](https://archive.is/ucKET) [[ 1 b]](https://web.archive.org/web/20260512034507/https://professional.bloomberg.com/products/data/enterprise-catalog/event-driven-feeds/) | | Source types | Open-web and institutional text [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Bloomberg reporting plus selected third-party, web, and social content [[ 2 a]](https://archive.is/rarjM) [[ 2 b]](https://web.archive.org/web/20260713190534/https://assets.bbhub.io/professional/sites/41/Fact-Sheet-EDF-Textual-News.pdf) | | Published source metric | 300,000+ open-web sources [[ 4 a]](https://archive.is/glqrw) | Cited 2015 fact sheet: more than 100,000 sources [[ 2 a]](https://archive.is/rarjM) [[ 2 b]](https://web.archive.org/web/20260713190534/https://assets.bbhub.io/professional/sites/41/Fact-Sheet-EDF-Textual-News.pdf) | | History | About 30 years [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Cited 2015 fact sheet: textual-news archive dating to 1992 [[ 2 a]](https://archive.is/rarjM) [[ 2 b]](https://web.archive.org/web/20260713190534/https://assets.bbhub.io/professional/sites/41/Fact-Sheet-EDF-Textual-News.pdf) | | Point-in-time method | Five-way date verification [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Real-time structured delivery [[ 1 a]](https://archive.is/ucKET) [[ 1 b]](https://web.archive.org/web/20260512034507/https://professional.bloomberg.com/products/data/enterprise-catalog/event-driven-feeds/) | | Events | 100M+ ranked dated events [[ 5 a]](https://archive.is/tGYc9) | Corporate-event calendar and market-event data [[ 1 a]](https://archive.is/ucKET) [[ 1 b]](https://web.archive.org/web/20260512034507/https://professional.bloomberg.com/products/data/enterprise-catalog/event-driven-feeds/) | | Delivery | API, SDKs, MCP, Cybernaut-1 [[ 4 a]](https://archive.is/glqrw) | Machine-readable enterprise feeds [[ 1 a]](https://archive.is/ucKET) [[ 1 b]](https://web.archive.org/web/20260512034507/https://professional.bloomberg.com/products/data/enterprise-catalog/event-driven-feeds/) | | Pricing | Product access through NOSIBLE [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Contact Bloomberg for access and commercial terms [[ 1 a]](https://archive.is/ucKET) [[ 1 b]](https://web.archive.org/web/20260512034507/https://professional.bloomberg.com/products/data/enterprise-catalog/event-driven-feeds/) | ## Potential fit by workflow Bloomberg publishes Event-Driven Feeds for enterprise market-data and systematic-feed infrastructure. [[ 1 a]](https://archive.is/ucKET) [[ 1 b]](https://web.archive.org/web/20260512034507/https://professional.bloomberg.com/products/data/enterprise-catalog/event-driven-feeds/) NOSIBLE addresses a different layer: dated government, company, regional, and specialist web evidence that agents can retrieve and inspect. [[ 4 a]](https://archive.is/glqrw) In NOSIBLE's view, the products may be complementary when a team needs both normalized feeds and open-web source discovery. ## Common Bloomberg comparison questions ### How does NOSIBLE feel it differentiates itself from Bloomberg? NOSIBLE is an AI-native company with two products: SEARCH and WORLD. [[ 5 a]](https://archive.is/tGYc9) [[ 7 a]](https://archive.is/iDZeo) SEARCH lets agents find dated open-web sources they can cite and inspect directly. [[ 4 a]](https://archive.is/glqrw) WORLD is a live open-web event database for models and backtests, with an embedding per event. [[ 5 a]](https://archive.is/tGYc9) [[ 6 a]](https://archive.is/tnkpG) [[ 8 a]](https://archive.is/1CUvQ) NOSIBLE is committed to open-source software and makes its models publicly available on Hugging Face. [[ 9 a]](https://archive.is/7SSz0) [[ 10 a]](https://archive.is/kHxMG) Related [WORLD v1.2 trial](https://nosible.com/start-trial#data-coverage) [Sentiment model](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) [Forward-looking model](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) ### Can NOSIBLE replace Bloomberg Event-Driven Feeds for systematic trading? NOSIBLE is not a terminal or a substitute for Bloomberg's low-latency market-data infrastructure. [[ 1 a]](https://archive.is/ucKET) [[ 1 b]](https://web.archive.org/web/20260512034507/https://professional.bloomberg.com/products/data/enterprise-catalog/event-driven-feeds/) [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) It is for teams that need multilingual source discovery, ranked open-web events, and replayable evidence. [[ 1 a]](https://archive.is/ucKET) [[ 1 b]](https://web.archive.org/web/20260512034507/https://professional.bloomberg.com/products/data/enterprise-catalog/event-driven-feeds/) [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) The evaluation should separate feed latency and entitlements from source breadth, historical retrieval, and inspectability. Related [WORLD event database](https://nosible.world/world) [Stock Screening](https://nosible.com/#proof) ### How does NOSIBLE's point-in-time method differ from Bloomberg's real-time feed model? The cited Bloomberg materials describe real-time structured delivery and a textual-news archive. [[ 1 a]](https://archive.is/ucKET) [[ 1 b]](https://web.archive.org/web/20260512034507/https://professional.bloomberg.com/products/data/enterprise-catalog/event-driven-feeds/) NOSIBLE's cited materials describe date verification for open-web documents and events. [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) These are different temporal models. [[ 1 a]](https://archive.is/ucKET) [[ 1 b]](https://web.archive.org/web/20260512034507/https://professional.bloomberg.com/products/data/enterprise-catalog/event-driven-feeds/) [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Teams should test each against the exact information set used by a model. Related [WORLD event database](https://nosible.world/world) ### What if we already use Bloomberg Terminal, Data License, or B-PIPE? In NOSIBLE's view, one possible division is to retain Bloomberg for licensed news, identifiers, real-time feeds, and other market-data workflows while using NOSIBLE for open-web retrieval and event evidence. Buyers should test overlap and incremental coverage rather than assume that either dataset contains every relevant source. Related [Copyright posture](https://nosible.com/legal/copyright) [WORLD event database](https://nosible.world/world) [Bulk Web Search](https://docs.nosible.com/endpoints/search-bulk) ### How does NOSIBLE compare with Bloomberg's tickerized real-time news feeds? NOSIBLE connects open-web events to tickers and other entities while retaining the original source context. [[ 5 a]](https://archive.is/tGYc9) Bloomberg's feed also attaches company, topic, and people metadata to textual news. [[ 1 a]](https://archive.is/ucKET) [[ 1 b]](https://web.archive.org/web/20260512034507/https://professional.bloomberg.com/products/data/enterprise-catalog/event-driven-feeds/) Teams should compare mapping precision and source coverage on their own universe rather than rely on headline counts. Related [WORLD event database](https://nosible.world/world) ### Do both products organize information around the same object? Bloomberg's cited feed materials describe company, topic, people, and other metadata attached to textual news. [[ 1 a]](https://archive.is/ucKET) [[ 1 b]](https://web.archive.org/web/20260512034507/https://professional.bloomberg.com/products/data/enterprise-catalog/event-driven-feeds/) NOSIBLE starts from documents and events, then associates the evidence with tickers, organizations, people, places, products, and risk concepts. [[ 5 a]](https://archive.is/tGYc9) [[ 6 a]](https://archive.is/tnkpG) In NOSIBLE's view, some teams may need both layers. Related [WORLD event database](https://nosible.world/world) Continue comparing ## Related Vendor Comparisons Compare Bloomberg with adjacent options for task execution, agent tooling, data access, and research workflows, then evaluate the output shape your application actually needs for production. [RavenPack](https://nosible.com/compare/nosible-vs-ravenpack) [Dow Jones](https://nosible.com/compare/nosible-vs-dow-jones) [LSEG](https://nosible.com/compare/nosible-vs-lseg) [Kensho](https://nosible.com/compare/nosible-vs-kensho) Dive deeper ## Take your next step today Review the delivered field definitions, classification boundaries, and example values before comparing vendor workflows, so your team can assess what each product actually returns to downstream systems. [Data dictionaries](https://nosible.com/data-dictionaries) [Ontology reference](https://nosible.com/ontologies) [API reference](https://docs.nosible.com/) [Start trial](https://nosible.com/start-trial) Sources reviewed July 14, 2026 : [[ 1 ] Bloomberg Event-Driven Feeds](https://professional.bloomberg.com/products/data/enterprise-catalog/event-driven-feeds/) ( [dated snapshot](https://archive.is/ucKET); [Wayback copy](https://web.archive.org/web/20260512034507/https://professional.bloomberg.com/products/data/enterprise-catalog/event-driven-feeds/) ) , [[ 2 ] Bloomberg Textual News fact sheet](https://assets.bbhub.io/professional/sites/41/Fact-Sheet-EDF-Textual-News.pdf) ( [dated snapshot](https://archive.is/rarjM); [Wayback copy](https://web.archive.org/web/20260713190534/https://assets.bbhub.io/professional/sites/41/Fact-Sheet-EDF-Textual-News.pdf) ) , [[ 3 ] NOSIBLE product overview](https://nosible.com/) ( [dated snapshot](https://archive.is/r7qDv); [Wayback copy](https://web.archive.org/web/20260713185114/https://nosible.com/) ) , [[ 4 ] NOSIBLE SEARCH](https://docs.nosible.com/) ( [dated snapshot](https://archive.is/glqrw) ) , [[ 5 ] NOSIBLE WORLD](https://nosible.world/world) ( [dated snapshot](https://archive.is/tGYc9) ) , [[ 6 ] NOSIBLE WORLD v1.2 trial and coverage](https://nosible.com/start-trial) ( [dated snapshot](https://archive.is/tnkpG) ) , [[ 7 ] NOSIBLE AI-native research overview](https://nosible.com/blog) ( [dated snapshot](https://archive.is/iDZeo) ) , [[ 8 ] NOSIBLE embedding-based research](https://nosible.com/blog/an-embedding-based-approach-to-trade-and-economic-policy-uncertainty) ( [dated snapshot](https://archive.is/1CUvQ) ) , [[ 9 ] NOSIBLE Financial Sentiment v1.2 Base](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) ( [dated snapshot](https://archive.is/7SSz0) ) , [[ 10 ] NOSIBLE Forward-Looking v1.2 Base](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) ( [dated snapshot](https://archive.is/kHxMG) ) . Bloomberg and Bloomberg Terminal are trademarks of Bloomberg Finance L.P. or its affiliates. NOSIBLE is not affiliated with, sponsored by, or endorsed by Bloomberg. Comparison standards, legal context & corrections This comparison was prepared by NOSIBLE, which has a commercial interest in the products being compared. It is based on the cited public materials as they appeared on July 14, 2026. NOSIBLE has not tested every competitor feature, and the page is not a complete statement of either product. Products change; confirm current requirements, availability, and commercial terms with each vendor. Factual claims are attributed to cited first-party materials; evaluative statements reflect NOSIBLE's opinion. Each source includes a dated archive.is snapshot and, where the Internet Archive captured that URL, a timestamp-specific Wayback copy. Archive availability is controlled by those services. Competitor names and marks are used only to identify the products being compared. No affiliation, sponsorship, or endorsement is implied. The FTC says truthful, non-deceptive comparative advertising may identify competitors ( [policy](https://www.ftc.gov/legal-library/browse/statement-policy-regarding-comparative-advertising); [dated snapshot](https://archive.is/nZXY4); [Wayback copy](https://web.archive.org/web/20260215042003/https://www.ftc.gov/legal-library/browse/statement-policy-regarding-comparative-advertising) ). For one U.S. example of nominative-use analysis, see *New Kids on the Block v. News America Publishing, Inc.*, 971 F.2d 302, 308 (9th Cir. 1992) ( [opinion](https://law.justia.com/cases/federal/appellate-courts/F2/971/302/72076/); [dated snapshot](https://archive.is/dnk8j); [Wayback copy](https://web.archive.org/web/20250211213616/https://law.justia.com/cases/federal/appellate-courts/F2/971/302/72076/) ). If you represent Bloomberg and believe a factual statement is inaccurate, email [stuart@nosible.com](mailto:stuart@nosible.com?subject=Correction%20request%3A%20NOSIBLE%20vs%20Bloomberg) with the specific claim and a supporting first-party URL. NOSIBLE will review and correct substantiated errors. > Compare NOSIBLE and Bloomberg for point-in-time event intelligence, open-web search, AI agents, backtesting, multilingual coverage, delivery, and market signal discovery. **URL:** https://nosible.com/compare/nosible-vs-bloomberg --- --- title: "NOSIBLE WORLD vs LSEG" description: "A NOSIBLE-authored comparison of NOSIBLE WORLD, a web-scale search and market-event intelligence engine for AI agents, and LSEG (formerly Refinitiv) Machine Readable News. Coverage, point-in-time integrity, identifiers, delivery, and pricing." url: "https://nosible.com/compare/nosible-vs-lseg" --- Comparison / Reviewed July 14, 2026 # NOSIBLE WORLD vs LSEG LSEG Machine Readable News turns Reuters and third-party news into a real-time feed with analytics and market identifiers. [[ 2 a]](https://archive.is/2zLU7) [[ 2 b]](https://web.archive.org/web/20260512091035/https://www.lseg.com/en/data-analytics/financial-data/financial-news-coverage/political-news-feeds-analysis/news-analytics) NOSIBLE WORLD focuses on dated open-web documents and ranked events for agents, research, and historical analysis. [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) NOSIBLE-AUTHORED COMPARISON · NOSIBLE HAS A COMMERCIAL INTEREST IN THIS COMPARISON · FACTS ATTRIBUTED TO FIRST-PARTY VENDOR MATERIALS · EVALUATIVE STATEMENTS ARE NOSIBLE'S OPINION · REVIEWED JULY 14, 2026 · DUAL PUBLIC ARCHIVES WHERE SUPPORTED · [STANDARDS & CORRECTIONS](https://nosible.com/compare/nosible-vs-lseg#comparison-standards). IF YOU REPRESENT LSEG AND BELIEVE A FACTUAL STATEMENT IS INACCURATE, EMAIL [STUART@NOSIBLE.COM](mailto:stuart@nosible.com?subject=Correction%20request%3A%20NOSIBLE%20WORLD%20vs%20LSEG) WITH THE SPECIFIC CLAIM AND A SUPPORTING FIRST-PARTY URL. NOSIBLE WILL REVIEW AND CORRECT SUBSTANTIATED ERRORS. - LSEG says it analyzes news across 35,000+ companies in real time; NOSIBLE WORLD v1.2 reports 38,097 stable ticker mappings. [[ 5 a]](https://archive.is/tGYc9) [[ 6 a]](https://archive.is/tnkpG) Company coverage and ticker mappings are different units. [[ 5 a]](https://archive.is/tGYc9) - Its News Analytics measures company sentiment, relevance, and novelty and supplies 90 metadata fields. [[ 2 a]](https://archive.is/2zLU7) [[ 2 b]](https://web.archive.org/web/20260512091035/https://www.lseg.com/en/data-analytics/financial-data/financial-news-coverage/political-news-feeds-analysis/news-analytics) - LSEG reports an analytics archive back to 2003, RIC and PermID tags, and delivery from sub-3-millisecond headlines to daily aggregates. [[ 1 a]](https://archive.is/pJnxL) [[ 1 b]](https://web.archive.org/web/20250606210418/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/machine-readable-news-and-quantitative-data-factsheet.pdf) - NOSIBLE WORLD emphasizes dated open-web retrieval, ranked events, multilingual evidence, and agent access. [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) ## Machine-readable news and source retrieval LSEG's published offering is a quantitative-news product with Reuters content, low-latency delivery, company and commodity analytics, historical scores, and market identifiers. [[ 2 a]](https://archive.is/2zLU7) [[ 2 b]](https://web.archive.org/web/20260512091035/https://www.lseg.com/en/data-analytics/financial-data/financial-news-coverage/political-news-feeds-analysis/news-analytics) NOSIBLE WORLD is a different kind of system, centered on dated open-web documents and ranked events that agents can query and inspect. [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) In NOSIBLE's view, NOSIBLE is the stronger fit when source breadth and replayable evidence are more important than wire latency. ## Feature comparison The competitor column summarizes what the named vendor's cited first-party materials describe; the NOSIBLE column summarizes NOSIBLE's cited materials. Category NOSIBLE Web-scale search and market-event intelligence for AI agents [[ 4 a]](https://archive.is/glqrw) LSEG Machine-readable Reuters and third-party news plus News Analytics [[ 2 a]](https://archive.is/2zLU7) [[ 2 b]](https://web.archive.org/web/20260512091035/https://www.lseg.com/en/data-analytics/financial-data/financial-news-coverage/political-news-feeds-analysis/news-analytics) Content scope NOSIBLE Long-form news, corporate, government; excludes social, paywalled, marketplaces, adult [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) LSEG Reuters News and third-party services [[ 2 a]](https://archive.is/2zLU7) [[ 2 b]](https://web.archive.org/web/20260512091035/https://www.lseg.com/en/data-analytics/financial-data/financial-news-coverage/political-news-feeds-analysis/news-analytics) History and point-in-time NOSIBLE About 30 years; five-way date-verification method [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) LSEG News Analytics archive with millisecond timestamps back to 2003 [[ 1 a]](https://archive.is/pJnxL) [[ 1 b]](https://web.archive.org/web/20250606210418/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/machine-readable-news-and-quantitative-data-factsheet.pdf) Published event-analysis scale NOSIBLE Standalone 100M+ dated, ranked events (WORLD) [[ 5 a]](https://archive.is/tGYc9) LSEG Cited fact sheet: real-time news analysis across 35,000+ companies [[ 1 a]](https://archive.is/pJnxL) [[ 1 b]](https://web.archive.org/web/20250606210418/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/machine-readable-news-and-quantitative-data-factsheet.pdf) Sentiment NOSIBLE Available in enrichment [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) LSEG Company sentiment, relevance, and novelty [[ 2 a]](https://archive.is/2zLU7) [[ 2 b]](https://web.archive.org/web/20260512091035/https://www.lseg.com/en/data-analytics/financial-data/financial-news-coverage/political-news-feeds-analysis/news-analytics) Analytics level NOSIBLE On the roadmap [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) LSEG Company- and commodity-level story scores [[ 1 a]](https://archive.is/pJnxL) [[ 1 b]](https://web.archive.org/web/20250606210418/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/machine-readable-news-and-quantitative-data-factsheet.pdf) Entity and identifier mapping NOSIBLE 3.2M organizations; 38,097 stable ticker mappings [[ 5 a]](https://archive.is/tGYc9) [[ 6 a]](https://archive.is/tnkpG) LSEG RIC identifiers and PermID tags [[ 1 a]](https://archive.is/pJnxL) [[ 1 b]](https://web.archive.org/web/20250606210418/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/machine-readable-news-and-quantitative-data-factsheet.pdf) Delivery and API NOSIBLE Six-endpoint agent API, Python and TypeScript SDKs, MCP [[ 4 a]](https://archive.is/glqrw) LSEG Latency spectrum from sub-3-millisecond headlines to one-day aggregates [[ 1 a]](https://archive.is/pJnxL) [[ 1 b]](https://web.archive.org/web/20250606210418/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/machine-readable-news-and-quantitative-data-factsheet.pdf) Machine-readable design NOSIBLE Built for AI agents [[ 4 a]](https://archive.is/glqrw) LSEG 90 metadata fields and machine-readable feeds [[ 1 a]](https://archive.is/pJnxL) [[ 1 b]](https://web.archive.org/web/20250606210418/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/machine-readable-news-and-quantitative-data-factsheet.pdf) Published use cases NOSIBLE Product materials describe event research and backtesting workflows [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) LSEG Published use cases include event trading, sentiment strategies, surveillance, and model training [[ 2 a]](https://archive.is/2zLU7) [[ 2 b]](https://web.archive.org/web/20260512091035/https://www.lseg.com/en/data-analytics/financial-data/financial-news-coverage/political-news-feeds-analysis/news-analytics) Pricing NOSIBLE See nosible.com [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) LSEG Contact LSEG for access and commercial terms [[ 1 a]](https://archive.is/pJnxL) [[ 1 b]](https://web.archive.org/web/20250606210418/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/machine-readable-news-and-quantitative-data-factsheet.pdf) | Dimension | NOSIBLE | LSEG | | --- | --- | --- | | Category | Web-scale search and market-event intelligence for AI agents [[ 4 a]](https://archive.is/glqrw) | Machine-readable Reuters and third-party news plus News Analytics [[ 2 a]](https://archive.is/2zLU7) [[ 2 b]](https://web.archive.org/web/20260512091035/https://www.lseg.com/en/data-analytics/financial-data/financial-news-coverage/political-news-feeds-analysis/news-analytics) | | Content scope | Long-form news, corporate, government; excludes social, paywalled, marketplaces, adult [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Reuters News and third-party services [[ 2 a]](https://archive.is/2zLU7) [[ 2 b]](https://web.archive.org/web/20260512091035/https://www.lseg.com/en/data-analytics/financial-data/financial-news-coverage/political-news-feeds-analysis/news-analytics) | | History and point-in-time | About 30 years; five-way date-verification method [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | News Analytics archive with millisecond timestamps back to 2003 [[ 1 a]](https://archive.is/pJnxL) [[ 1 b]](https://web.archive.org/web/20250606210418/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/machine-readable-news-and-quantitative-data-factsheet.pdf) | | Published event-analysis scale | Standalone 100M+ dated, ranked events (WORLD) [[ 5 a]](https://archive.is/tGYc9) | Cited fact sheet: real-time news analysis across 35,000+ companies [[ 1 a]](https://archive.is/pJnxL) [[ 1 b]](https://web.archive.org/web/20250606210418/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/machine-readable-news-and-quantitative-data-factsheet.pdf) | | Sentiment | Available in enrichment [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Company sentiment, relevance, and novelty [[ 2 a]](https://archive.is/2zLU7) [[ 2 b]](https://web.archive.org/web/20260512091035/https://www.lseg.com/en/data-analytics/financial-data/financial-news-coverage/political-news-feeds-analysis/news-analytics) | | Analytics level | On the roadmap [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Company- and commodity-level story scores [[ 1 a]](https://archive.is/pJnxL) [[ 1 b]](https://web.archive.org/web/20250606210418/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/machine-readable-news-and-quantitative-data-factsheet.pdf) | | Entity and identifier mapping | 3.2M organizations; 38,097 stable ticker mappings [[ 5 a]](https://archive.is/tGYc9) [[ 6 a]](https://archive.is/tnkpG) | RIC identifiers and PermID tags [[ 1 a]](https://archive.is/pJnxL) [[ 1 b]](https://web.archive.org/web/20250606210418/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/machine-readable-news-and-quantitative-data-factsheet.pdf) | | Delivery and API | Six-endpoint agent API, Python and TypeScript SDKs, MCP [[ 4 a]](https://archive.is/glqrw) | Latency spectrum from sub-3-millisecond headlines to one-day aggregates [[ 1 a]](https://archive.is/pJnxL) [[ 1 b]](https://web.archive.org/web/20250606210418/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/machine-readable-news-and-quantitative-data-factsheet.pdf) | | Machine-readable design | Built for AI agents [[ 4 a]](https://archive.is/glqrw) | 90 metadata fields and machine-readable feeds [[ 1 a]](https://archive.is/pJnxL) [[ 1 b]](https://web.archive.org/web/20250606210418/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/machine-readable-news-and-quantitative-data-factsheet.pdf) | | Published use cases | Product materials describe event research and backtesting workflows [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Published use cases include event trading, sentiment strategies, surveillance, and model training [[ 2 a]](https://archive.is/2zLU7) [[ 2 b]](https://web.archive.org/web/20260512091035/https://www.lseg.com/en/data-analytics/financial-data/financial-news-coverage/political-news-feeds-analysis/news-analytics) | | Pricing | See nosible.com [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Contact LSEG for access and commercial terms [[ 1 a]](https://archive.is/pJnxL) [[ 1 b]](https://web.archive.org/web/20250606210418/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/machine-readable-news-and-quantitative-data-factsheet.pdf) | ## Potential fit by workflow LSEG's cited materials describe Reuters content, low-latency delivery, news analytics, and RIC-based workflows. [[ 2 a]](https://archive.is/2zLU7) [[ 2 b]](https://web.archive.org/web/20260512091035/https://www.lseg.com/en/data-analytics/financial-data/financial-news-coverage/political-news-feeds-analysis/news-analytics) NOSIBLE's cited materials describe dated open-web source retrieval, ranked events, and historical evidence for agents. [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) In NOSIBLE's view, consider the product whose published workflow matches the requirement; a combined architecture may also be appropriate. ## Common LSEG comparison questions ### How does NOSIBLE feel it differentiates itself from LSEG? NOSIBLE is an AI-native company with two products: SEARCH and WORLD. [[ 5 a]](https://archive.is/tGYc9) [[ 7 a]](https://archive.is/iDZeo) SEARCH lets agents find dated open-web sources they can cite and inspect directly. [[ 4 a]](https://archive.is/glqrw) WORLD is a live open-web event database for models and backtests, with an embedding per event. [[ 5 a]](https://archive.is/tGYc9) [[ 6 a]](https://archive.is/tnkpG) [[ 8 a]](https://archive.is/1CUvQ) NOSIBLE is committed to open-source software and makes its models publicly available on Hugging Face. [[ 9 a]](https://archive.is/7SSz0) [[ 10 a]](https://archive.is/kHxMG) Related [WORLD v1.2 trial](https://nosible.com/start-trial#data-coverage) [Sentiment model](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) [Forward-looking model](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) ### Can NOSIBLE replace LSEG Machine Readable News if we need Reuters content? No. NOSIBLE does not provide a Reuters license. [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) If Reuters content is required, retain the appropriate LSEG product and rights. [[ 2 a]](https://archive.is/2zLU7) [[ 2 b]](https://web.archive.org/web/20260512091035/https://www.lseg.com/en/data-analytics/financial-data/financial-news-coverage/political-news-feeds-analysis/news-analytics) NOSIBLE can be evaluated separately as an open-web document-and-event layer with a different source universe. [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Related [WORLD event database](https://nosible.world/world) ### How does NOSIBLE compare with RIC and PermID-mapped News Analytics? LSEG publishes RIC identifiers and PermID tags for interoperability. [[ 1 a]](https://archive.is/pJnxL) [[ 1 b]](https://web.archive.org/web/20250606210418/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/machine-readable-news-and-quantitative-data-factsheet.pdf) NOSIBLE associates source evidence with tickers, organizations, people, places, products, and risk concepts. [[ 5 a]](https://archive.is/tGYc9) [[ 6 a]](https://archive.is/tnkpG) Buyers should map the two identifier systems against their own security master before assuming interchangeability. Related [WORLD event database](https://nosible.world/world) ### What if our strategy depends on ultra-low-latency headlines? LSEG's fact sheet advertises delivery from less than three milliseconds for an ultra-low-latency headline feed. [[ 1 a]](https://archive.is/pJnxL) [[ 1 b]](https://web.archive.org/web/20250606210418/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/machine-readable-news-and-quantitative-data-factsheet.pdf) In NOSIBLE's view, LSEG is the clearer fit when that delivery model is required. NOSIBLE should instead be evaluated for source retrieval, ranked events, agents, and historical evidence—not as an HFT wire replacement. [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) Related [WORLD event database](https://nosible.world/world) ### Is LSEG's historical News Analytics enough for backtesting? LSEG publishes millisecond-stamped News Analytics back to 2003 and expressly positions the archive for model training and backtesting. [[ 1 a]](https://archive.is/pJnxL) [[ 1 b]](https://web.archive.org/web/20250606210418/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/machine-readable-news-and-quantitative-data-factsheet.pdf) NOSIBLE targets a different open-web information set. [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) The right test is whether each historical query reproduces the documents, identifiers, and timestamps required by the strategy. Related [WORLD event database](https://nosible.world/world) ### How does NOSIBLE differ from LSEG's finished News Analytics scores? LSEG's cited materials describe company and commodity sentiment, relevance, and novelty scores. [[ 2 a]](https://archive.is/2zLU7) [[ 2 b]](https://web.archive.org/web/20260512091035/https://www.lseg.com/en/data-analytics/financial-data/financial-news-coverage/political-news-feeds-analysis/news-analytics) NOSIBLE provides dated documents and events from which teams can build or validate their own signals. [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Compare the desired output—finished analytics or source evidence—before choosing. Related [Sentiment model](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) [Sentiment research](https://nosible.com/blog/fast-enough-to-matter-productionizing-tiny-transformers-for-signal-extraction) [WORLD event database](https://nosible.world/world) Continue comparing ## Related Vendor Comparisons Compare LSEG with adjacent options for task execution, agent tooling, data access, and research workflows, then evaluate the output shape your application actually needs for production. [Bloomberg](https://nosible.com/compare/nosible-vs-bloomberg) [RavenPack](https://nosible.com/compare/nosible-vs-ravenpack) [MarketPsych](https://nosible.com/compare/nosible-vs-marketpsych) [Kensho](https://nosible.com/compare/nosible-vs-kensho) Dive deeper ## Take your next step today Review the delivered field definitions, classification boundaries, and example values before comparing vendor workflows, so your team can assess what each product actually returns to downstream systems. [Data dictionaries](https://nosible.com/data-dictionaries) [Ontology reference](https://nosible.com/ontologies) [API reference](https://docs.nosible.com/) [Start trial](https://nosible.com/start-trial) Sources reviewed July 14, 2026 : [[ 1 ] LSEG Machine Readable News fact sheet](https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/machine-readable-news-and-quantitative-data-factsheet.pdf) ( [dated snapshot](https://archive.is/pJnxL); [Wayback copy](https://web.archive.org/web/20250606210418/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/machine-readable-news-and-quantitative-data-factsheet.pdf) ) , [[ 2 ] LSEG News Analytics](https://www.lseg.com/en/data-analytics/financial-data/financial-news-coverage/political-news-feeds-analysis/news-analytics) ( [dated snapshot](https://archive.is/2zLU7); [Wayback copy](https://web.archive.org/web/20260512091035/https://www.lseg.com/en/data-analytics/financial-data/financial-news-coverage/political-news-feeds-analysis/news-analytics) ) , [[ 3 ] NOSIBLE product overview](https://nosible.com/) ( [dated snapshot](https://archive.is/r7qDv); [Wayback copy](https://web.archive.org/web/20260713185114/https://nosible.com/) ) , [[ 4 ] NOSIBLE SEARCH](https://docs.nosible.com/) ( [dated snapshot](https://archive.is/glqrw) ) , [[ 5 ] NOSIBLE WORLD](https://nosible.world/world) ( [dated snapshot](https://archive.is/tGYc9) ) , [[ 6 ] NOSIBLE WORLD v1.2 trial and coverage](https://nosible.com/start-trial) ( [dated snapshot](https://archive.is/tnkpG) ) , [[ 7 ] NOSIBLE AI-native research overview](https://nosible.com/blog) ( [dated snapshot](https://archive.is/iDZeo) ) , [[ 8 ] NOSIBLE embedding-based research](https://nosible.com/blog/an-embedding-based-approach-to-trade-and-economic-policy-uncertainty) ( [dated snapshot](https://archive.is/1CUvQ) ) , [[ 9 ] NOSIBLE Financial Sentiment v1.2 Base](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) ( [dated snapshot](https://archive.is/7SSz0) ) , [[ 10 ] NOSIBLE Forward-Looking v1.2 Base](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) ( [dated snapshot](https://archive.is/kHxMG) ) . This page compares NOSIBLE WORLD with LSEG Machine Readable News and News Analytics, not LSEG as a whole. LSEG, Refinitiv, Reuters, RIC, and PermID are marks of their respective owners. NOSIBLE is not affiliated with, sponsored by, or endorsed by LSEG. Comparison standards, legal context & corrections This comparison was prepared by NOSIBLE, which has a commercial interest in the products being compared. It is based on the cited public materials as they appeared on July 14, 2026. NOSIBLE has not tested every competitor feature, and the page is not a complete statement of either product. Products change; confirm current requirements, availability, and commercial terms with each vendor. Factual claims are attributed to cited first-party materials; evaluative statements reflect NOSIBLE's opinion. Each source includes a dated archive.is snapshot and, where the Internet Archive captured that URL, a timestamp-specific Wayback copy. Archive availability is controlled by those services. Competitor names and marks are used only to identify the products being compared. No affiliation, sponsorship, or endorsement is implied. The FTC says truthful, non-deceptive comparative advertising may identify competitors ( [policy](https://www.ftc.gov/legal-library/browse/statement-policy-regarding-comparative-advertising); [dated snapshot](https://archive.is/nZXY4); [Wayback copy](https://web.archive.org/web/20260215042003/https://www.ftc.gov/legal-library/browse/statement-policy-regarding-comparative-advertising) ). For one U.S. example of nominative-use analysis, see *New Kids on the Block v. News America Publishing, Inc.*, 971 F.2d 302, 308 (9th Cir. 1992) ( [opinion](https://law.justia.com/cases/federal/appellate-courts/F2/971/302/72076/); [dated snapshot](https://archive.is/dnk8j); [Wayback copy](https://web.archive.org/web/20250211213616/https://law.justia.com/cases/federal/appellate-courts/F2/971/302/72076/) ). If you represent LSEG and believe a factual statement is inaccurate, email [stuart@nosible.com](mailto:stuart@nosible.com?subject=Correction%20request%3A%20NOSIBLE%20WORLD%20vs%20LSEG) with the specific claim and a supporting first-party URL. NOSIBLE will review and correct substantiated errors. > A NOSIBLE-authored comparison of NOSIBLE WORLD, a web-scale search and market-event intelligence engine for AI agents, and LSEG (formerly Refinitiv) Machine Readable News. Coverage, point-in-time integrity, identifiers, delivery, and pricing. **URL:** https://nosible.com/compare/nosible-vs-lseg --- --- title: "NOSIBLE vs Dow Jones" description: "Compare NOSIBLE and Dow Jones for open-web event intelligence, point-in-time research, AI agents, language coverage, source access, delivery, and backtesting." url: "https://nosible.com/compare/nosible-vs-dow-jones" --- Comparison / Reviewed July 14, 2026 # NOSIBLE vs Dow Jones Dow Jones Factiva combines licensed global news, company data, monitoring, feeds, and APIs. [[ 2 a]](https://archive.is/elL66) NOSIBLE focuses on dated open-web retrieval and ranked events for agents and research. [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) NOSIBLE-AUTHORED COMPARISON · NOSIBLE HAS A COMMERCIAL INTEREST IN THIS COMPARISON · FACTS ATTRIBUTED TO FIRST-PARTY VENDOR MATERIALS · EVALUATIVE STATEMENTS ARE NOSIBLE'S OPINION · REVIEWED JULY 14, 2026 · DUAL PUBLIC ARCHIVES WHERE SUPPORTED · [STANDARDS & CORRECTIONS](https://nosible.com/compare/nosible-vs-dow-jones#comparison-standards). IF YOU REPRESENT DOW JONES AND BELIEVE A FACTUAL STATEMENT IS INACCURATE, EMAIL [STUART@NOSIBLE.COM](mailto:stuart@nosible.com?subject=Correction%20request%3A%20NOSIBLE%20vs%20Dow%20Jones) WITH THE SPECIFIC CLAIM AND A SUPPORTING FIRST-PARTY URL. NOSIBLE WILL REVIEW AND CORRECT SUBSTANTIATED ERRORS. - Dow Jones says Factiva covers 33,000 sources across 200 countries and 32 languages. [[ 1 a]](https://archive.is/xeOGz) [[ 1 b]](https://web.archive.org/web/20260710105426/https://www.dowjones.com/business-intelligence/factiva/) - Factiva reports 35 million company profiles; NOSIBLE WORLD v1.2 reports about 3.2 million organizations and 6.4 million people. [[ 5 a]](https://archive.is/tGYc9) [[ 6 a]](https://archive.is/tnkpG) Profiles and extracted entities are not equivalent records. [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) - Dow Jones offers licensed content through monitoring products, feeds, APIs, and GenAI-oriented solutions. [[ 2 a]](https://archive.is/elL66) - NOSIBLE emphasizes dated source retrieval, ranked events, multilingual open-web coverage, and agent access. [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) - In NOSIBLE's view, consider NOSIBLE when source-level event discovery matters more than a licensed news catalog. ## Licensed intelligence and open-web evidence Factiva is a licensed-news and business-intelligence product with monitoring, company information, feeds, APIs, and content prepared for GenAI use. [[ 2 a]](https://archive.is/elL66) NOSIBLE is designed around dated open-web sources and ranked events that agents and research systems can inspect directly. [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) In NOSIBLE's view, NOSIBLE is the stronger fit when a team wants to construct its own event workflow from source evidence. ## Licensed news and dated event evidence The competitor column summarizes what the named vendor's cited first-party materials describe; the NOSIBLE column summarizes NOSIBLE's cited materials. Primary use NOSIBLE Point-in-time event search for AI agents [[ 4 a]](https://archive.is/glqrw) Dow Jones Licensed global news and business intelligence [[ 1 a]](https://archive.is/xeOGz) [[ 1 b]](https://web.archive.org/web/20260710105426/https://www.dowjones.com/business-intelligence/factiva/) Illustrative users NOSIBLE AI agents, backtests, risk systems [[ 4 a]](https://archive.is/glqrw) Dow Jones Research, strategy, communications, and risk teams [[ 1 a]](https://archive.is/xeOGz) [[ 1 b]](https://web.archive.org/web/20260710105426/https://www.dowjones.com/business-intelligence/factiva/) Content scope NOSIBLE Long-form news, corporate, and government text [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Dow Jones 33,000 premium news and data sources [[ 1 a]](https://archive.is/xeOGz) [[ 1 b]](https://web.archive.org/web/20260710105426/https://www.dowjones.com/business-intelligence/factiva/) Source model NOSIBLE Open-web and institutional text [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Dow Jones Licensed content across 200 countries [[ 1 a]](https://archive.is/xeOGz) [[ 1 b]](https://web.archive.org/web/20260710105426/https://www.dowjones.com/business-intelligence/factiva/) Languages NOSIBLE 95 [[ 4 a]](https://archive.is/glqrw) Dow Jones 32 [[ 1 a]](https://archive.is/xeOGz) [[ 1 b]](https://web.archive.org/web/20260710105426/https://www.dowjones.com/business-intelligence/factiva/) Point-in-time method NOSIBLE Five-way date verification [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Dow Jones Live monitoring plus feed and retrieval access [[ 1 a]](https://archive.is/xeOGz) [[ 1 b]](https://web.archive.org/web/20260710105426/https://www.dowjones.com/business-intelligence/factiva/) Event model and workflows NOSIBLE 100M+ ranked dated events [[ 5 a]](https://archive.is/tGYc9) Dow Jones News monitoring, trends, and business-intelligence workflows [[ 1 a]](https://archive.is/xeOGz) [[ 1 b]](https://web.archive.org/web/20260710105426/https://www.dowjones.com/business-intelligence/factiva/) Delivery NOSIBLE API, SDKs, MCP, Cybernaut-1 [[ 4 a]](https://archive.is/glqrw) Dow Jones Factiva platform, feeds, APIs, and developer tools [[ 2 a]](https://archive.is/elL66) Pricing NOSIBLE Product access through NOSIBLE [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Dow Jones Contact Dow Jones for access and commercial terms [[ 1 a]](https://archive.is/xeOGz) [[ 1 b]](https://web.archive.org/web/20260710105426/https://www.dowjones.com/business-intelligence/factiva/) | Dimension | NOSIBLE | Dow Jones | | --- | --- | --- | | Primary use | Point-in-time event search for AI agents [[ 4 a]](https://archive.is/glqrw) | Licensed global news and business intelligence [[ 1 a]](https://archive.is/xeOGz) [[ 1 b]](https://web.archive.org/web/20260710105426/https://www.dowjones.com/business-intelligence/factiva/) | | Illustrative users | AI agents, backtests, risk systems [[ 4 a]](https://archive.is/glqrw) | Research, strategy, communications, and risk teams [[ 1 a]](https://archive.is/xeOGz) [[ 1 b]](https://web.archive.org/web/20260710105426/https://www.dowjones.com/business-intelligence/factiva/) | | Content scope | Long-form news, corporate, and government text [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | 33,000 premium news and data sources [[ 1 a]](https://archive.is/xeOGz) [[ 1 b]](https://web.archive.org/web/20260710105426/https://www.dowjones.com/business-intelligence/factiva/) | | Source model | Open-web and institutional text [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Licensed content across 200 countries [[ 1 a]](https://archive.is/xeOGz) [[ 1 b]](https://web.archive.org/web/20260710105426/https://www.dowjones.com/business-intelligence/factiva/) | | Languages | 95 [[ 4 a]](https://archive.is/glqrw) | 32 [[ 1 a]](https://archive.is/xeOGz) [[ 1 b]](https://web.archive.org/web/20260710105426/https://www.dowjones.com/business-intelligence/factiva/) | | Point-in-time method | Five-way date verification [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Live monitoring plus feed and retrieval access [[ 1 a]](https://archive.is/xeOGz) [[ 1 b]](https://web.archive.org/web/20260710105426/https://www.dowjones.com/business-intelligence/factiva/) | | Event model and workflows | 100M+ ranked dated events [[ 5 a]](https://archive.is/tGYc9) | News monitoring, trends, and business-intelligence workflows [[ 1 a]](https://archive.is/xeOGz) [[ 1 b]](https://web.archive.org/web/20260710105426/https://www.dowjones.com/business-intelligence/factiva/) | | Delivery | API, SDKs, MCP, Cybernaut-1 [[ 4 a]](https://archive.is/glqrw) | Factiva platform, feeds, APIs, and developer tools [[ 2 a]](https://archive.is/elL66) | | Pricing | Product access through NOSIBLE [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Contact Dow Jones for access and commercial terms [[ 1 a]](https://archive.is/xeOGz) [[ 1 b]](https://web.archive.org/web/20260710105426/https://www.dowjones.com/business-intelligence/factiva/) | ## Potential fit by workflow Dow Jones positions Factiva around its licensed corpus, monitoring tools, company profiles, and delivery options. [[ 1 a]](https://archive.is/xeOGz) [[ 1 b]](https://web.archive.org/web/20260710105426/https://www.dowjones.com/business-intelligence/factiva/) NOSIBLE is aimed at the adjacent source layer: dated government, company, regional, and specialist web evidence connected to events. [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) In NOSIBLE's view, buyers may retain licensed journalism while adding open-web discovery where it is useful. ## Common Dow Jones comparison questions ### How does NOSIBLE feel it differentiates itself from Dow Jones? NOSIBLE is an AI-native company with two products: SEARCH and WORLD. [[ 5 a]](https://archive.is/tGYc9) [[ 7 a]](https://archive.is/iDZeo) SEARCH lets agents find dated open-web sources they can cite and inspect directly. [[ 4 a]](https://archive.is/glqrw) WORLD is a live open-web event database for models and backtests, with an embedding per event. [[ 5 a]](https://archive.is/tGYc9) [[ 6 a]](https://archive.is/tnkpG) [[ 8 a]](https://archive.is/1CUvQ) NOSIBLE is committed to open-source software and makes its models publicly available on Hugging Face. [[ 9 a]](https://archive.is/7SSz0) [[ 10 a]](https://archive.is/kHxMG) Related [WORLD v1.2 trial](https://nosible.com/start-trial#data-coverage) [Sentiment model](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) [Forward-looking model](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) ### Can NOSIBLE replace Factiva or Dow Jones Newswires content? No. NOSIBLE does not grant rights to Dow Jones-owned or third-party licensed journalism. [[ 1 a]](https://archive.is/xeOGz) [[ 1 b]](https://web.archive.org/web/20260710105426/https://www.dowjones.com/business-intelligence/factiva/) [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Teams that require Factiva content should retain the appropriate Dow Jones license and evaluate NOSIBLE separately for incremental open-web evidence. [[ 1 a]](https://archive.is/xeOGz) [[ 1 b]](https://web.archive.org/web/20260710105426/https://www.dowjones.com/business-intelligence/factiva/) [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Related [WORLD event database](https://nosible.world/world) ### How does NOSIBLE differ from Dow Jones DNA snapshots and streams? Factiva Feeds and APIs are the direct Dow Jones option for integrating licensed news and data into applications. [[ 2 a]](https://archive.is/elL66) NOSIBLE offers a different corpus and event model. [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) The practical question is whether the workflow requires Dow Jones content rights, open-web event discovery, or both. Related [WORLD event database](https://nosible.world/world) ### How does NOSIBLE compare with Factiva Sentiment Signals? NOSIBLE is not presented as a replacement for every packaged indicator available through Dow Jones. [[ 1 a]](https://archive.is/xeOGz) [[ 1 b]](https://web.archive.org/web/20260710105426/https://www.dowjones.com/business-intelligence/factiva/) [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Its value is access to dated sources, ranked events, entity mappings, and enrichment paths that teams can use to construct custom signals. [[ 1 a]](https://archive.is/xeOGz) [[ 1 b]](https://web.archive.org/web/20260710105426/https://www.dowjones.com/business-intelligence/factiva/) [[ 5 a]](https://archive.is/tGYc9) Compare outputs on a defined historical sample. Related [Sentiment model](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) [Sentiment research](https://nosible.com/blog/fast-enough-to-matter-productionizing-tiny-transformers-for-signal-extraction) [WORLD event database](https://nosible.world/world) ### What should buyers ask about licensing for AI, RAG, or model training? Ask both vendors about retrieval, embeddings, model training, storage, excerpts, and redistribution. Dow Jones expressly markets Factiva AI Research Solutions as licensed content for GenAI use, while NOSIBLE is designed around agentic retrieval and source attribution. [[ 1 a]](https://archive.is/xeOGz) [[ 1 b]](https://web.archive.org/web/20260710105426/https://www.dowjones.com/business-intelligence/factiva/) [[ 4 a]](https://archive.is/glqrw) Contract terms, not this comparison, determine permitted use. Related [Copyright posture](https://nosible.com/legal/copyright) [Agentic Search](https://docs.nosible.com/) ### What if our strategy depends on low-latency market-news distribution? Keep the wire or feed when its latency, editorial coverage, and licensing are required. Evaluate NOSIBLE for agentic research, ranked open-web events, and historical source reconstruction. A combined test should measure incremental coverage and timing without assuming that either service is universally earlier. Related [Agentic Search](https://docs.nosible.com/) [WORLD event database](https://nosible.world/world) Continue comparing ## Related Vendor Comparisons Compare Dow Jones with adjacent options for task execution, agent tooling, data access, and research workflows, then evaluate the output shape your application actually needs for production. [Bloomberg](https://nosible.com/compare/nosible-vs-bloomberg) [AlphaSense](https://nosible.com/compare/nosible-vs-alphasense) [RavenPack](https://nosible.com/compare/nosible-vs-ravenpack) [Opoint](https://nosible.com/compare/nosible-vs-opoint) Dive deeper ## Take your next step today Review the delivered field definitions, classification boundaries, and example values before comparing vendor workflows, so your team can assess what each product actually returns to downstream systems. [Data dictionaries](https://nosible.com/data-dictionaries) [Ontology reference](https://nosible.com/ontologies) [API reference](https://docs.nosible.com/) [Start trial](https://nosible.com/start-trial) Sources reviewed July 14, 2026 : [[ 1 ] Dow Jones Factiva](https://www.dowjones.com/business-intelligence/factiva/) ( [dated snapshot](https://archive.is/xeOGz); [Wayback copy](https://web.archive.org/web/20260710105426/https://www.dowjones.com/business-intelligence/factiva/) ) , [[ 2 ] Dow Jones Factiva APIs](https://developer.dowjones.com/site/global/apis/factiva_analytics/index.gsp) ( [dated snapshot](https://archive.is/elL66) ) , [[ 3 ] NOSIBLE product overview](https://nosible.com/) ( [dated snapshot](https://archive.is/r7qDv); [Wayback copy](https://web.archive.org/web/20260713185114/https://nosible.com/) ) , [[ 4 ] NOSIBLE SEARCH](https://docs.nosible.com/) ( [dated snapshot](https://archive.is/glqrw) ) , [[ 5 ] NOSIBLE WORLD](https://nosible.world/world) ( [dated snapshot](https://archive.is/tGYc9) ) , [[ 6 ] NOSIBLE WORLD v1.2 trial and coverage](https://nosible.com/start-trial) ( [dated snapshot](https://archive.is/tnkpG) ) , [[ 7 ] NOSIBLE AI-native research overview](https://nosible.com/blog) ( [dated snapshot](https://archive.is/iDZeo) ) , [[ 8 ] NOSIBLE embedding-based research](https://nosible.com/blog/an-embedding-based-approach-to-trade-and-economic-policy-uncertainty) ( [dated snapshot](https://archive.is/1CUvQ) ) , [[ 9 ] NOSIBLE Financial Sentiment v1.2 Base](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) ( [dated snapshot](https://archive.is/7SSz0) ) , [[ 10 ] NOSIBLE Forward-Looking v1.2 Base](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) ( [dated snapshot](https://archive.is/kHxMG) ) . Dow Jones and Factiva are trademarks of Dow Jones & Company, Inc. or its affiliates. NOSIBLE is not affiliated with, sponsored by, or endorsed by Dow Jones. Comparison standards, legal context & corrections This comparison was prepared by NOSIBLE, which has a commercial interest in the products being compared. It is based on the cited public materials as they appeared on July 14, 2026. NOSIBLE has not tested every competitor feature, and the page is not a complete statement of either product. Products change; confirm current requirements, availability, and commercial terms with each vendor. Factual claims are attributed to cited first-party materials; evaluative statements reflect NOSIBLE's opinion. Each source includes a dated archive.is snapshot and, where the Internet Archive captured that URL, a timestamp-specific Wayback copy. Archive availability is controlled by those services. Competitor names and marks are used only to identify the products being compared. No affiliation, sponsorship, or endorsement is implied. The FTC says truthful, non-deceptive comparative advertising may identify competitors ( [policy](https://www.ftc.gov/legal-library/browse/statement-policy-regarding-comparative-advertising); [dated snapshot](https://archive.is/nZXY4); [Wayback copy](https://web.archive.org/web/20260215042003/https://www.ftc.gov/legal-library/browse/statement-policy-regarding-comparative-advertising) ). For one U.S. example of nominative-use analysis, see *New Kids on the Block v. News America Publishing, Inc.*, 971 F.2d 302, 308 (9th Cir. 1992) ( [opinion](https://law.justia.com/cases/federal/appellate-courts/F2/971/302/72076/); [dated snapshot](https://archive.is/dnk8j); [Wayback copy](https://web.archive.org/web/20250211213616/https://law.justia.com/cases/federal/appellate-courts/F2/971/302/72076/) ). If you represent Dow Jones and believe a factual statement is inaccurate, email [stuart@nosible.com](mailto:stuart@nosible.com?subject=Correction%20request%3A%20NOSIBLE%20vs%20Dow%20Jones) with the specific claim and a supporting first-party URL. NOSIBLE will review and correct substantiated errors. > Compare NOSIBLE and Dow Jones for open-web event intelligence, point-in-time research, AI agents, language coverage, source access, delivery, and backtesting. **URL:** https://nosible.com/compare/nosible-vs-dow-jones --- --- title: "NOSIBLE vs AlphaSense" description: "Compare NOSIBLE and AlphaSense for market research, source access, AI workflows, point-in-time analysis, monitoring, and enterprise intelligence." url: "https://nosible.com/compare/nosible-vs-alphasense" --- Comparison / Reviewed July 25, 2026 # NOSIBLE vs AlphaSense AlphaSense describes a premium research workspace for market intelligence. [[ 1 a]](https://archive.is/VY5F6) [[ 1 b]](https://web.archive.org/web/20260725105019/https://www.alpha-sense.com/platform/) [[ 2 a]](https://archive.is/6ybAF) [[ 2 b]](https://web.archive.org/web/20260725105126/https://www.alpha-sense.com/solutions/market-intelligence-platform/) NOSIBLE gives research agents dated open-web sources they can inspect, plus ranked events for time-bound analysis. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) NOSIBLE-AUTHORED COMPARISON · NOSIBLE HAS A COMMERCIAL INTEREST IN THIS COMPARISON · FACTS ATTRIBUTED TO FIRST-PARTY VENDOR MATERIALS · EVALUATIVE STATEMENTS ARE NOSIBLE'S OPINION · REVIEWED JULY 25, 2026 · DUAL PUBLIC ARCHIVES WHERE SUPPORTED · [STANDARDS & CORRECTIONS](https://nosible.com/compare/nosible-vs-alphasense#comparison-standards). IF YOU REPRESENT ALPHASENSE AND BELIEVE A FACTUAL STATEMENT IS INACCURATE, EMAIL [STUART@NOSIBLE.COM](mailto:stuart@nosible.com?subject=Correction%20request%3A%20NOSIBLE%20vs%20AlphaSense) WITH THE SPECIFIC CLAIM AND A SUPPORTING FIRST-PARTY URL. NOSIBLE WILL REVIEW AND CORRECT SUBSTANTIATED ERRORS. - AlphaSense describes one platform spanning Generative Search, Deep Research, Enterprise Intelligence, financial data, workflow agents, and monitoring. [[ 1 a]](https://archive.is/VY5F6) [[ 1 b]](https://web.archive.org/web/20260725105019/https://www.alpha-sense.com/platform/) - AlphaSense says its platform combines internal content with private, public, premium, and proprietary external materials. [[ 1 a]](https://archive.is/VY5F6) [[ 1 b]](https://web.archive.org/web/20260725105019/https://www.alpha-sense.com/platform/) [[ 2 a]](https://archive.is/6ybAF) [[ 2 b]](https://web.archive.org/web/20260725105126/https://www.alpha-sense.com/solutions/market-intelligence-platform/) - NOSIBLE SEARCH gives agents dated sources they can cite and inspect; WORLD adds ranked events with an embedding per event, rather than a premium research library. [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) [[ 8 a]](https://archive.is/1CUvQ) - In NOSIBLE's view, AlphaSense may be the more direct fit for teams that require its premium, proprietary, or internal-content workflow. - In NOSIBLE's view, NOSIBLE is the stronger fit when inspectable open-web evidence and replayable event history are central to the task. ## The research workspace versus the evidence trail AlphaSense describes a market-intelligence platform that brings search, deep research, financial data, monitoring, and customer content together. [[ 1 a]](https://archive.is/VY5F6) [[ 1 b]](https://web.archive.org/web/20260725105019/https://www.alpha-sense.com/platform/) NOSIBLE begins where a researcher needs to inspect dated open-web sources or follow a ranked event through time. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 5 a]](https://archive.is/fv0cj) In NOSIBLE's view, that makes the products complementary when a premium research workspace still needs a transparent external evidence trail. ## Premium research workspace and inspectable evidence The competitor column summarizes what the named vendor's cited first-party materials describe; the NOSIBLE column summarizes NOSIBLE's cited materials. Primary use NOSIBLE Dated open-web intelligence for agents and research [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 4 a]](https://archive.is/YwGWR) AlphaSense Market intelligence and enterprise research workflows [[ 1 a]](https://archive.is/VY5F6) [[ 1 b]](https://web.archive.org/web/20260725105019/https://www.alpha-sense.com/platform/) Research inputs NOSIBLE Open-web sources and ranked events [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 5 a]](https://archive.is/fv0cj) AlphaSense Premium, proprietary, public, private, and internal content [[ 2 a]](https://archive.is/6ybAF) [[ 2 b]](https://web.archive.org/web/20260725105126/https://www.alpha-sense.com/solutions/market-intelligence-platform/) Search workflow NOSIBLE SEARCH API with agent, SDK, and MCP access [[ 4 a]](https://archive.is/YwGWR) AlphaSense Generative Search and Deep Research [[ 1 a]](https://archive.is/VY5F6) [[ 1 b]](https://web.archive.org/web/20260725105019/https://www.alpha-sense.com/platform/) Enterprise content NOSIBLE Open-web retrieval [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) AlphaSense AlphaSense describes Enterprise Intelligence for customer content [[ 1 a]](https://archive.is/VY5F6) [[ 1 b]](https://web.archive.org/web/20260725105019/https://www.alpha-sense.com/platform/) Monitoring NOSIBLE Dated source and event retrieval [[ 4 a]](https://archive.is/YwGWR) AlphaSense Monitoring is part of AlphaSense's published platform [[ 1 a]](https://archive.is/VY5F6) [[ 1 b]](https://web.archive.org/web/20260725105019/https://www.alpha-sense.com/platform/) Point-in-time method NOSIBLE Documented buyer evaluation AlphaSense Evaluate historical-research behavior with a buyer test Event representation NOSIBLE Ranked dated WORLD events [[ 5 a]](https://archive.is/fv0cj) AlphaSense Research outputs and financial-data workflows [[ 1 a]](https://archive.is/VY5F6) [[ 1 b]](https://web.archive.org/web/20260725105019/https://www.alpha-sense.com/platform/) Delivery NOSIBLE API, SDKs, MCP, and web products [[ 4 a]](https://archive.is/YwGWR) AlphaSense Platform and developer integration options [[ 1 a]](https://archive.is/VY5F6) [[ 1 b]](https://web.archive.org/web/20260725105019/https://www.alpha-sense.com/platform/) Pricing NOSIBLE Confirm current terms with NOSIBLE AlphaSense Confirm current terms with AlphaSense | Dimension | NOSIBLE | AlphaSense | | --- | --- | --- | | Primary use | Dated open-web intelligence for agents and research [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 4 a]](https://archive.is/YwGWR) | Market intelligence and enterprise research workflows [[ 1 a]](https://archive.is/VY5F6) [[ 1 b]](https://web.archive.org/web/20260725105019/https://www.alpha-sense.com/platform/) | | Research inputs | Open-web sources and ranked events [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 5 a]](https://archive.is/fv0cj) | Premium, proprietary, public, private, and internal content [[ 2 a]](https://archive.is/6ybAF) [[ 2 b]](https://web.archive.org/web/20260725105126/https://www.alpha-sense.com/solutions/market-intelligence-platform/) | | Search workflow | SEARCH API with agent, SDK, and MCP access [[ 4 a]](https://archive.is/YwGWR) | Generative Search and Deep Research [[ 1 a]](https://archive.is/VY5F6) [[ 1 b]](https://web.archive.org/web/20260725105019/https://www.alpha-sense.com/platform/) | | Enterprise content | Open-web retrieval [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | AlphaSense describes Enterprise Intelligence for customer content [[ 1 a]](https://archive.is/VY5F6) [[ 1 b]](https://web.archive.org/web/20260725105019/https://www.alpha-sense.com/platform/) | | Monitoring | Dated source and event retrieval [[ 4 a]](https://archive.is/YwGWR) | Monitoring is part of AlphaSense's published platform [[ 1 a]](https://archive.is/VY5F6) [[ 1 b]](https://web.archive.org/web/20260725105019/https://www.alpha-sense.com/platform/) | | Point-in-time method | Documented buyer evaluation | Evaluate historical-research behavior with a buyer test | | Event representation | Ranked dated WORLD events [[ 5 a]](https://archive.is/fv0cj) | Research outputs and financial-data workflows [[ 1 a]](https://archive.is/VY5F6) [[ 1 b]](https://web.archive.org/web/20260725105019/https://www.alpha-sense.com/platform/) | | Delivery | API, SDKs, MCP, and web products [[ 4 a]](https://archive.is/YwGWR) | Platform and developer integration options [[ 1 a]](https://archive.is/VY5F6) [[ 1 b]](https://web.archive.org/web/20260725105019/https://www.alpha-sense.com/platform/) | | Pricing | Confirm current terms with NOSIBLE | Confirm current terms with AlphaSense | ## When premium research matters—and when the trail matters In NOSIBLE's view, AlphaSense may be the more direct choice for a team that needs premium, proprietary, financial-data, or internal-content workflows in one workspace. NOSIBLE is designed for a different moment in the process: agents can work from dated sources and ranked events, with an embedding per event for downstream analysis. [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) [[ 8 a]](https://archive.is/1CUvQ) In NOSIBLE's view, that is a meaningful complement rather than a substitute claim. ## Common AlphaSense comparison questions ### How does NOSIBLE feel it differentiates itself from AlphaSense? NOSIBLE is an AI-native company with two products: SEARCH and WORLD. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 5 a]](https://archive.is/fv0cj) [[ 7 a]](https://archive.is/iDZeo) SEARCH lets agents find dated open-web sources they can cite and inspect directly. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 4 a]](https://archive.is/YwGWR) WORLD is a live open-web event database for models and backtests, with an embedding per event. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 5 a]](https://archive.is/fv0cj) [[ 6 a]](https://archive.is/tnkpG) [[ 8 a]](https://archive.is/1CUvQ) NOSIBLE is committed to open-source software and makes its models publicly available on Hugging Face. [[ 9 a]](https://archive.is/7SSz0) [[ 10 a]](https://archive.is/kHxMG) Related [WORLD v1.2 trial](https://nosible.com/start-trial#data-coverage) [Sentiment model](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) [Forward-looking model](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) ### Can NOSIBLE replace AlphaSense's premium and internal-content workflow? AlphaSense says its platform combines premium external material and customers' internal content with AI search, research, and monitoring. [[ 1 a]](https://archive.is/VY5F6) [[ 1 b]](https://web.archive.org/web/20260725105019/https://www.alpha-sense.com/platform/) [[ 2 a]](https://archive.is/6ybAF) [[ 2 b]](https://web.archive.org/web/20260725105126/https://www.alpha-sense.com/solutions/market-intelligence-platform/) This page does not present NOSIBLE as a substitute for those collections. In NOSIBLE's view, the products can be complementary when an analyst needs proprietary research alongside inspectable open-web evidence. Related [WORLD event database](https://nosible.world/world) ### How do the products differ for AI research workflows? AlphaSense describes Generative Search and Deep Research within its market-intelligence platform. [[ 1 a]](https://archive.is/VY5F6) [[ 1 b]](https://web.archive.org/web/20260725105019/https://www.alpha-sense.com/platform/) NOSIBLE SEARCH is designed for agents to retrieve dated open-web sources, while WORLD supplies ranked events for models and backtests. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) In NOSIBLE's view, buyers should test each product against their required sources, citations, delivery method, and review process. Related [WORLD event database](https://nosible.world/world) ### Does AlphaSense offer a more direct fit for financial research? AlphaSense publishes financial-data tools, premium research content, and market-intelligence workflows for professional users. [[ 1 a]](https://archive.is/VY5F6) [[ 1 b]](https://web.archive.org/web/20260725105019/https://www.alpha-sense.com/platform/) [[ 2 a]](https://archive.is/6ybAF) [[ 2 b]](https://web.archive.org/web/20260725105126/https://www.alpha-sense.com/solutions/market-intelligence-platform/) NOSIBLE publishes open-web retrieval and market-event products. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) In NOSIBLE's view, AlphaSense may be the more direct fit when its proprietary financial-data and integrated-workspace features are required. Related [WORLD event database](https://nosible.world/world) ### How should a team compare historical research and monitoring? Ask each vendor what a historical query returns, how source revisions are handled, and which dates are exposed to users. AlphaSense publishes monitoring and research tools, while NOSIBLE emphasizes dated sources and ranked events. [[ 1 a]](https://archive.is/VY5F6) [[ 1 b]](https://web.archive.org/web/20260725105019/https://www.alpha-sense.com/platform/) [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) A fixed, documented test set is the most reliable way to compare coverage, timing, and citation needs. Related [WORLD event database](https://nosible.world/world) ### Can the two products be used together? Yes. AlphaSense describes a platform for premium, proprietary, and internal research material, while NOSIBLE focuses on dated open-web source retrieval and event history. [[ 1 a]](https://archive.is/VY5F6) [[ 1 b]](https://web.archive.org/web/20260725105019/https://www.alpha-sense.com/platform/) [[ 2 a]](https://archive.is/6ybAF) [[ 2 b]](https://web.archive.org/web/20260725105126/https://www.alpha-sense.com/solutions/market-intelligence-platform/) [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) In NOSIBLE's view, a team can retain AlphaSense for its enterprise workflow while using NOSIBLE to locate, inspect, and replay complementary open-web evidence. Related [WORLD event database](https://nosible.world/world) Continue comparing ## Related Vendor Comparisons Compare AlphaSense with adjacent options for task execution, agent tooling, data access, and research workflows, then evaluate the output shape your application actually needs for production. [Kensho](https://nosible.com/compare/nosible-vs-kensho) [Bloomberg](https://nosible.com/compare/nosible-vs-bloomberg) [Dow Jones](https://nosible.com/compare/nosible-vs-dow-jones) [Opoint](https://nosible.com/compare/nosible-vs-opoint) Dive deeper ## Take your next step today Review the delivered field definitions, classification boundaries, and example values before comparing vendor workflows, so your team can assess what each product actually returns to downstream systems. [Data dictionaries](https://nosible.com/data-dictionaries) [Ontology reference](https://nosible.com/ontologies) [API reference](https://docs.nosible.com/) [Start trial](https://nosible.com/start-trial) Sources reviewed July 25, 2026 : [[ 1 ] AlphaSense platform](https://www.alpha-sense.com/platform/) ( [dated snapshot](https://archive.is/VY5F6); [Wayback copy](https://web.archive.org/web/20260725105019/https://www.alpha-sense.com/platform/) ) , [[ 2 ] AlphaSense market intelligence platform](https://www.alpha-sense.com/solutions/market-intelligence-platform/) ( [dated snapshot](https://archive.is/6ybAF); [Wayback copy](https://web.archive.org/web/20260725105126/https://www.alpha-sense.com/solutions/market-intelligence-platform/) ) , [[ 3 ] NOSIBLE product overview](https://nosible.com/) ( [dated snapshot](https://archive.is/8abI7); [Wayback copy](https://web.archive.org/web/20260713185114/https://nosible.com/) ) , [[ 4 ] NOSIBLE SEARCH](https://docs.nosible.com/) ( [dated snapshot](https://archive.is/YwGWR) ) , [[ 5 ] NOSIBLE WORLD](https://nosible.world/world) ( [dated snapshot](https://archive.is/fv0cj) ) , [[ 6 ] NOSIBLE WORLD v1.2 trial and coverage](https://nosible.com/start-trial) ( [dated snapshot](https://archive.is/tnkpG) ) , [[ 7 ] NOSIBLE AI-native research overview](https://nosible.com/blog) ( [dated snapshot](https://archive.is/iDZeo) ) , [[ 8 ] NOSIBLE embedding-based research](https://nosible.com/blog/an-embedding-based-approach-to-trade-and-economic-policy-uncertainty) ( [dated snapshot](https://archive.is/1CUvQ) ) , [[ 9 ] NOSIBLE Financial Sentiment v1.2 Base](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) ( [dated snapshot](https://archive.is/7SSz0) ) , [[ 10 ] NOSIBLE Forward-Looking v1.2 Base](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) ( [dated snapshot](https://archive.is/kHxMG) ) . AlphaSense is used solely to identify the compared product. NOSIBLE is not affiliated with, sponsored by, or endorsed by AlphaSense. Comparison standards, legal context & corrections This comparison was prepared by NOSIBLE, which has a commercial interest in the products being compared. It is based on the cited public materials as they appeared on July 25, 2026. NOSIBLE has not tested every competitor feature, and the page is not a complete statement of either product. Products change; confirm current requirements, availability, and commercial terms with each vendor. Factual claims are attributed to cited first-party materials; evaluative statements reflect NOSIBLE's opinion. Each source includes a dated archive.is snapshot and, where the Internet Archive captured that URL, a timestamp-specific Wayback copy. Archive availability is controlled by those services. Competitor names and marks are used only to identify the products being compared. No affiliation, sponsorship, or endorsement is implied. The FTC says truthful, non-deceptive comparative advertising may identify competitors ( [policy](https://www.ftc.gov/legal-library/browse/statement-policy-regarding-comparative-advertising); [dated snapshot](https://archive.is/nZXY4); [Wayback copy](https://web.archive.org/web/20260215042003/https://www.ftc.gov/legal-library/browse/statement-policy-regarding-comparative-advertising) ). For one U.S. example of nominative-use analysis, see *New Kids on the Block v. News America Publishing, Inc.*, 971 F.2d 302, 308 (9th Cir. 1992) ( [opinion](https://law.justia.com/cases/federal/appellate-courts/F2/971/302/72076/); [dated snapshot](https://archive.is/dnk8j); [Wayback copy](https://web.archive.org/web/20250211213616/https://law.justia.com/cases/federal/appellate-courts/F2/971/302/72076/) ). If you represent AlphaSense and believe a factual statement is inaccurate, email [stuart@nosible.com](mailto:stuart@nosible.com?subject=Correction%20request%3A%20NOSIBLE%20vs%20AlphaSense) with the specific claim and a supporting first-party URL. NOSIBLE will review and correct substantiated errors. > Compare NOSIBLE and AlphaSense for market research, source access, AI workflows, point-in-time analysis, monitoring, and enterprise intelligence. **URL:** https://nosible.com/compare/nosible-vs-alphasense --- --- title: "NOSIBLE vs Kensho" description: "Compare NOSIBLE and Kensho for financial AI data APIs, S&P data access, cited retrieval, open-web evidence, ranked events, and research-agent workflows." url: "https://nosible.com/compare/nosible-vs-kensho" --- Comparison / Reviewed July 25, 2026 # NOSIBLE vs Kensho Kensho documents AI-ready retrieval across S&P data. [[ 1 a]](https://archive.is/aG4al) [[ 2 a]](https://archive.is/CpIA5) NOSIBLE is the complementary open-web evidence and ranked-event layer for teams whose research questions extend beyond an entitled data universe. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) NOSIBLE-AUTHORED COMPARISON · NOSIBLE HAS A COMMERCIAL INTEREST IN THIS COMPARISON · FACTS ATTRIBUTED TO FIRST-PARTY VENDOR MATERIALS · EVALUATIVE STATEMENTS ARE NOSIBLE'S OPINION · REVIEWED JULY 25, 2026 · DUAL PUBLIC ARCHIVES WHERE SUPPORTED · [STANDARDS & CORRECTIONS](https://nosible.com/compare/nosible-vs-kensho#comparison-standards). IF YOU REPRESENT KENSHO AND BELIEVE A FACTUAL STATEMENT IS INACCURATE, EMAIL [STUART@NOSIBLE.COM](mailto:stuart@nosible.com?subject=Correction%20request%3A%20NOSIBLE%20vs%20Kensho) WITH THE SPECIFIC CLAIM AND A SUPPORTING FIRST-PARTY URL. NOSIBLE WILL REVIEW AND CORRECT SUBSTANTIATED ERRORS. - Kensho documents an LLM-ready API for querying S&P data through Python and MCP. [[ 1 a]](https://archive.is/aG4al) [[ 2 a]](https://archive.is/CpIA5) - Kensho says Adaptive Retrieval uses one endpoint across S&P data and returns source information and citations. [[ 1 a]](https://archive.is/aG4al) [[ 2 a]](https://archive.is/CpIA5) - NOSIBLE SEARCH and WORLD focus on dated open-web sources and ranked events, with an embedding per event, rather than S&P data access. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 5 a]](https://archive.is/fv0cj) [[ 8 a]](https://archive.is/1CUvQ) - In NOSIBLE's view, Kensho may be the more direct fit when a workflow requires its S&P data and cited retrieval interface. - In NOSIBLE's view, NOSIBLE is the stronger fit when open-web source evidence and event-oriented historical analysis are central. ## Start with the data entitlement, not the interface Kensho documents LLM-ready and adaptive retrieval APIs for querying S&P data, including source information and citations. [[ 1 a]](https://archive.is/aG4al) [[ 2 a]](https://archive.is/CpIA5) NOSIBLE supplies dated open-web sources and ranked events. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 5 a]](https://archive.is/fv0cj) In NOSIBLE's view, the first procurement question is data domain: a team that requires S&P datasets should assess Kensho directly; a team that needs inspectable open-web evidence and event history should assess NOSIBLE directly. ## Entitled financial data and open-web event evidence The competitor column summarizes what the named vendor's cited first-party materials describe; the NOSIBLE column summarizes NOSIBLE's cited materials. Primary use NOSIBLE Dated open-web source and event intelligence [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Kensho AI-ready retrieval across S&P data [[ 1 a]](https://archive.is/aG4al) [[ 2 a]](https://archive.is/CpIA5) Core data NOSIBLE Open-web sources and ranked events [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 5 a]](https://archive.is/fv0cj) Kensho S&P data exposed through Kensho APIs [[ 1 a]](https://archive.is/aG4al) [[ 2 a]](https://archive.is/CpIA5) LLM interface NOSIBLE Agent API, SDKs, and MCP [[ 4 a]](https://archive.is/YwGWR) Kensho LLM-ready API with Python and MCP [[ 1 a]](https://archive.is/aG4al) Retrieval NOSIBLE Dated source retrieval [[ 4 a]](https://archive.is/YwGWR) Kensho Adaptive Retrieval through one endpoint across S&P data [[ 1 a]](https://archive.is/aG4al) [[ 2 a]](https://archive.is/CpIA5) Citations NOSIBLE Source-attributed evidence and event context Kensho Source information and citations in Adaptive Retrieval responses [[ 2 a]](https://archive.is/CpIA5) Point-in-time method NOSIBLE Documented buyer evaluation Kensho Evaluate entitlements and historical behavior with a buyer test Event representation NOSIBLE Ranked dated WORLD events [[ 5 a]](https://archive.is/fv0cj) Kensho Financial data retrieval rather than a published WORLD-style event database Delivery NOSIBLE API, SDKs, MCP, SEARCH, and WORLD [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) Kensho Kensho API interfaces Pricing NOSIBLE Confirm current terms with NOSIBLE Kensho Confirm current terms with Kensho and S&P Global | Dimension | NOSIBLE | Kensho | | --- | --- | --- | | Primary use | Dated open-web source and event intelligence [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | AI-ready retrieval across S&P data [[ 1 a]](https://archive.is/aG4al) [[ 2 a]](https://archive.is/CpIA5) | | Core data | Open-web sources and ranked events [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 5 a]](https://archive.is/fv0cj) | S&P data exposed through Kensho APIs [[ 1 a]](https://archive.is/aG4al) [[ 2 a]](https://archive.is/CpIA5) | | LLM interface | Agent API, SDKs, and MCP [[ 4 a]](https://archive.is/YwGWR) | LLM-ready API with Python and MCP [[ 1 a]](https://archive.is/aG4al) | | Retrieval | Dated source retrieval [[ 4 a]](https://archive.is/YwGWR) | Adaptive Retrieval through one endpoint across S&P data [[ 1 a]](https://archive.is/aG4al) [[ 2 a]](https://archive.is/CpIA5) | | Citations | Source-attributed evidence and event context | Source information and citations in Adaptive Retrieval responses [[ 2 a]](https://archive.is/CpIA5) | | Point-in-time method | Documented buyer evaluation | Evaluate entitlements and historical behavior with a buyer test | | Event representation | Ranked dated WORLD events [[ 5 a]](https://archive.is/fv0cj) | Financial data retrieval rather than a published WORLD-style event database | | Delivery | API, SDKs, MCP, SEARCH, and WORLD [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) | Kensho API interfaces | | Pricing | Confirm current terms with NOSIBLE | Confirm current terms with Kensho and S&P Global | ## When a closed data universe is the point—and when it is not In NOSIBLE's view, Kensho may be the more direct choice where S&P data, its access terms, and cited retrieval interface are non-negotiable. NOSIBLE is designed for dated open-web sources and ranked events, with an embedding per event for downstream analysis. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 5 a]](https://archive.is/fv0cj) [[ 8 a]](https://archive.is/1CUvQ) In NOSIBLE's view, the products become complementary only after the buyer validates entitlements, provenance, and the research role of each data type. ## Common Kensho comparison questions ### How does NOSIBLE feel it differentiates itself from Kensho? NOSIBLE is an AI-native company with two products: SEARCH and WORLD. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 5 a]](https://archive.is/fv0cj) [[ 7 a]](https://archive.is/iDZeo) SEARCH lets agents find dated open-web sources they can cite and inspect directly. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 4 a]](https://archive.is/YwGWR) WORLD is a live open-web event database for models and backtests, with an embedding per event. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 5 a]](https://archive.is/fv0cj) [[ 6 a]](https://archive.is/tnkpG) [[ 8 a]](https://archive.is/1CUvQ) NOSIBLE is committed to open-source software and makes its models publicly available on Hugging Face. [[ 9 a]](https://archive.is/7SSz0) [[ 10 a]](https://archive.is/kHxMG) Related [WORLD v1.2 trial](https://nosible.com/start-trial#data-coverage) [Sentiment model](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) [Forward-looking model](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) ### What does Kensho's LLM-ready API provide? Kensho documents an LLM-ready API for querying S&P data through Python and MCP. [[ 1 a]](https://archive.is/aG4al) [[ 2 a]](https://archive.is/CpIA5) NOSIBLE publishes agent access to dated open-web sources and ranked events. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) In NOSIBLE's view, Kensho may be the more direct fit when the required data is within its S&P offering and the application needs that retrieval interface. Related [Agentic Search](https://docs.nosible.com/) [WORLD event database](https://nosible.world/world) ### What does Kensho say about Adaptive Retrieval? Kensho says Adaptive Retrieval uses one endpoint across S&P data and returns source information and citations. [[ 1 a]](https://archive.is/aG4al) [[ 2 a]](https://archive.is/CpIA5) NOSIBLE emphasizes source-attributed open-web evidence and ranked events. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 5 a]](https://archive.is/fv0cj) In NOSIBLE's view, buyers should validate source scope, entitlement, citation detail, and historical availability using the records and access terms relevant to their workflow. Related [Agentic Search](https://docs.nosible.com/) [WORLD event database](https://nosible.world/world) ### Can NOSIBLE replace Kensho's S&P data access? Kensho's cited documentation concerns retrieval across S&P data, while NOSIBLE focuses on open-web sources and events. [[ 1 a]](https://archive.is/aG4al) [[ 2 a]](https://archive.is/CpIA5) [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) This page does not present NOSIBLE as a substitute for S&P data access. In NOSIBLE's view, a buyer should choose Kensho when that specific cited data access is a non-negotiable requirement. Related [WORLD event database](https://nosible.world/world) ### How should a financial team compare the products? Start with the required data rights, datasets, source types, timestamps, citations, and integration surface. [[ 2 a]](https://archive.is/CpIA5) Kensho publishes S&P-data retrieval, while NOSIBLE publishes dated open-web evidence and events. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) A documented evaluation should use representative securities and questions, then review access conditions and output provenance before any production deployment. Related [WORLD event database](https://nosible.world/world) [Bulk Web Search](https://docs.nosible.com/endpoints/search-bulk) ### Can Kensho and NOSIBLE be used together? Potentially. Kensho can serve S&P data through its published AI retrieval interfaces, while NOSIBLE can supply open-web sources and ranked event context. [[ 1 a]](https://archive.is/aG4al) [[ 2 a]](https://archive.is/CpIA5) [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 5 a]](https://archive.is/fv0cj) In NOSIBLE's view, teams should preserve provenance and ensure their licensing, historical use, and model-governance controls permit the intended combined workflow. Related [WORLD event database](https://nosible.world/world) Continue comparing ## Related Vendor Comparisons Compare Kensho with adjacent options for task execution, agent tooling, data access, and research workflows, then evaluate the output shape your application actually needs for production. [Bloomberg](https://nosible.com/compare/nosible-vs-bloomberg) [LSEG](https://nosible.com/compare/nosible-vs-lseg) [AlphaSense](https://nosible.com/compare/nosible-vs-alphasense) [RavenPack](https://nosible.com/compare/nosible-vs-ravenpack) Dive deeper ## Take your next step today Review the delivered field definitions, classification boundaries, and example values before comparing vendor workflows, so your team can assess what each product actually returns to downstream systems. [Data dictionaries](https://nosible.com/data-dictionaries) [Ontology reference](https://nosible.com/ontologies) [API reference](https://docs.nosible.com/) [Start trial](https://nosible.com/start-trial) Sources reviewed July 25, 2026 : [[ 1 ] Kensho LLM-ready API overview](https://docs.kensho.com/llmreadyapi/overview) ( [dated snapshot](https://archive.is/aG4al) ) , [[ 2 ] Kensho Adaptive Retrieval API guide](https://docs.kensho.com/adaptive-retrieval/api-guide) ( [dated snapshot](https://archive.is/CpIA5) ) , [[ 3 ] NOSIBLE product overview](https://nosible.com/) ( [dated snapshot](https://archive.is/8abI7); [Wayback copy](https://web.archive.org/web/20260713185114/https://nosible.com/) ) , [[ 4 ] NOSIBLE SEARCH](https://docs.nosible.com/) ( [dated snapshot](https://archive.is/YwGWR) ) , [[ 5 ] NOSIBLE WORLD](https://nosible.world/world) ( [dated snapshot](https://archive.is/fv0cj) ) , [[ 6 ] NOSIBLE WORLD v1.2 trial and coverage](https://nosible.com/start-trial) ( [dated snapshot](https://archive.is/tnkpG) ) , [[ 7 ] NOSIBLE AI-native research overview](https://nosible.com/blog) ( [dated snapshot](https://archive.is/iDZeo) ) , [[ 8 ] NOSIBLE embedding-based research](https://nosible.com/blog/an-embedding-based-approach-to-trade-and-economic-policy-uncertainty) ( [dated snapshot](https://archive.is/1CUvQ) ) , [[ 9 ] NOSIBLE Financial Sentiment v1.2 Base](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) ( [dated snapshot](https://archive.is/7SSz0) ) , [[ 10 ] NOSIBLE Forward-Looking v1.2 Base](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) ( [dated snapshot](https://archive.is/kHxMG) ) . Kensho and S&P Global are used solely to identify the products and data sources discussed. NOSIBLE is not affiliated with, sponsored by, or endorsed by Kensho or S&P Global. Comparison standards, legal context & corrections This comparison was prepared by NOSIBLE, which has a commercial interest in the products being compared. It is based on the cited public materials as they appeared on July 25, 2026. NOSIBLE has not tested every competitor feature, and the page is not a complete statement of either product. Products change; confirm current requirements, availability, and commercial terms with each vendor. Factual claims are attributed to cited first-party materials; evaluative statements reflect NOSIBLE's opinion. Each source includes a dated archive.is snapshot and, where the Internet Archive captured that URL, a timestamp-specific Wayback copy. Archive availability is controlled by those services. Competitor names and marks are used only to identify the products being compared. No affiliation, sponsorship, or endorsement is implied. The FTC says truthful, non-deceptive comparative advertising may identify competitors ( [policy](https://www.ftc.gov/legal-library/browse/statement-policy-regarding-comparative-advertising); [dated snapshot](https://archive.is/nZXY4); [Wayback copy](https://web.archive.org/web/20260215042003/https://www.ftc.gov/legal-library/browse/statement-policy-regarding-comparative-advertising) ). For one U.S. example of nominative-use analysis, see *New Kids on the Block v. News America Publishing, Inc.*, 971 F.2d 302, 308 (9th Cir. 1992) ( [opinion](https://law.justia.com/cases/federal/appellate-courts/F2/971/302/72076/); [dated snapshot](https://archive.is/dnk8j); [Wayback copy](https://web.archive.org/web/20250211213616/https://law.justia.com/cases/federal/appellate-courts/F2/971/302/72076/) ). If you represent Kensho and believe a factual statement is inaccurate, email [stuart@nosible.com](mailto:stuart@nosible.com?subject=Correction%20request%3A%20NOSIBLE%20vs%20Kensho) with the specific claim and a supporting first-party URL. NOSIBLE will review and correct substantiated errors. > Compare NOSIBLE and Kensho for financial AI data APIs, S&P data access, cited retrieval, open-web evidence, ranked events, and research-agent workflows. **URL:** https://nosible.com/compare/nosible-vs-kensho --- --- title: "NOSIBLE WORLD vs FactSet" description: "How NOSIBLE WORLD, a web-scale search and market-event intelligence engine for AI agents, compares with FactSet Truvalue Labs, an outside-in ESG signals provider. Coverage, point-in-time integrity, scope, and use cases." url: "https://nosible.com/compare/nosible-vs-factset-truvalue" --- Comparison / Reviewed July 14, 2026 # NOSIBLE WORLD vs FactSet FactSet Truvalue turns external ESG reporting into continuously updated company scores. [[ 1 a]](https://archive.is/eZ2mt) [[ 1 b]](https://web.archive.org/web/20221028182449/https://insight.factset.com/resources/at-a-glance-factset-truvalue-sasb-scores-datafeed) NOSIBLE WORLD is a domain-agnostic dated-document and event engine for agents, research, and custom risk models. [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) NOSIBLE-AUTHORED COMPARISON · NOSIBLE HAS A COMMERCIAL INTEREST IN THIS COMPARISON · FACTS ATTRIBUTED TO FIRST-PARTY VENDOR MATERIALS · EVALUATIVE STATEMENTS ARE NOSIBLE'S OPINION · REVIEWED JULY 14, 2026 · DUAL PUBLIC ARCHIVES WHERE SUPPORTED · [STANDARDS & CORRECTIONS](https://nosible.com/compare/nosible-vs-factset-truvalue#comparison-standards). IF YOU REPRESENT FACTSET AND BELIEVE A FACTUAL STATEMENT IS INACCURATE, EMAIL [STUART@NOSIBLE.COM](mailto:stuart@nosible.com?subject=Correction%20request%3A%20NOSIBLE%20WORLD%20vs%20FactSet) WITH THE SPECIFIC CLAIM AND A SUPPORTING FIRST-PARTY URL. NOSIBLE WILL REVIEW AND CORRECT SUBSTANTIATED ERRORS. - FactSet says Truvalue processes millions of documents from more than 150,000 sources in over 30 languages. [[ 1 a]](https://archive.is/eZ2mt) [[ 1 b]](https://web.archive.org/web/20221028182449/https://insight.factset.com/resources/at-a-glance-factset-truvalue-sasb-scores-datafeed) - Its cited sources include news, NGOs, watchdog groups, trade publications, and social media. [[ 1 a]](https://archive.is/eZ2mt) [[ 1 b]](https://web.archive.org/web/20221028182449/https://insight.factset.com/resources/at-a-glance-factset-truvalue-sasb-scores-datafeed) - Truvalue supplies ESG scores for 260,000+ companies; NOSIBLE WORLD v1.2 reports about 3.2 million organizations and 38,097 ticker mappings, but does not ship equivalent Truvalue scores. [[ 5 a]](https://archive.is/tGYc9) [[ 6 a]](https://archive.is/tnkpG) - FactSet's Truvalue brochure describes positive and negative ESG events and multiple score products. [[ 2 a]](https://archive.is/PK6Ir) [[ 2 b]](https://web.archive.org/web/20260713190535/https://advantage.factset.com/hubfs/Website/Resources%20Section/Brochures/esg-data-and-analytics-from-truvalue-labs-brochure.pdf) - The cited NOSIBLE materials describe a cross-domain event scope and emphasize dated sources, ranked events, and agent access. [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) ## ESG scores and cross-domain evidence FactSet Truvalue is an outside-in ESG dataset with defined categories, company coverage, and ready-made scores. [[ 1 a]](https://archive.is/eZ2mt) [[ 1 b]](https://web.archive.org/web/20221028182449/https://insight.factset.com/resources/at-a-glance-factset-truvalue-sasb-scores-datafeed) NOSIBLE WORLD is cross-domain: agents can retrieve dated sources and ranked events across ESG, policy, operations, litigation, supply chains, and other topics. [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) In NOSIBLE's view, NOSIBLE is the stronger fit when a team wants to build proprietary models from evidence rather than consume a finished ESG score. ## Feature comparison The competitor column summarizes what the named vendor's cited first-party materials describe; the NOSIBLE column summarizes NOSIBLE's cited materials. Category NOSIBLE Web-scale search and market-event intelligence for AI agents [[ 4 a]](https://archive.is/glqrw) FactSet Outside-in ESG signals and scores [[ 1 a]](https://archive.is/eZ2mt) [[ 1 b]](https://web.archive.org/web/20221028182449/https://insight.factset.com/resources/at-a-glance-factset-truvalue-sasb-scores-datafeed) Content scope NOSIBLE Long-form news, corporate, government; excludes social, paywalled, marketplaces, adult [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) FactSet News, trade journals, NGO and industry reports (third-party, not self-reported) [[ 1 a]](https://archive.is/eZ2mt) [[ 1 b]](https://web.archive.org/web/20221028182449/https://insight.factset.com/resources/at-a-glance-factset-truvalue-sasb-scores-datafeed) Languages NOSIBLE 95 [[ 4 a]](https://archive.is/glqrw) FactSet 30+ [[ 1 a]](https://archive.is/eZ2mt) [[ 1 b]](https://web.archive.org/web/20221028182449/https://insight.factset.com/resources/at-a-glance-factset-truvalue-sasb-scores-datafeed) History and point-in-time NOSIBLE About 30 years; five-way date-verification method [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) FactSet Data available back to 2007; event-driven data delivered daily [[ 1 a]](https://archive.is/eZ2mt) [[ 1 b]](https://web.archive.org/web/20221028182449/https://insight.factset.com/resources/at-a-glance-factset-truvalue-sasb-scores-datafeed) Event database NOSIBLE Standalone 100M+ dated, ranked events (WORLD) [[ 5 a]](https://archive.is/tGYc9) FactSet Positive and negative ESG events categorized by SASB topics [[ 2 a]](https://archive.is/PK6Ir) [[ 2 b]](https://web.archive.org/web/20260713190535/https://advantage.factset.com/hubfs/Website/Resources%20Section/Brochures/esg-data-and-analytics-from-truvalue-labs-brochure.pdf) Sentiment NOSIBLE Available in enrichment [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) FactSet Continuously updated material ESG scores [[ 2 a]](https://archive.is/PK6Ir) [[ 2 b]](https://web.archive.org/web/20260713190535/https://advantage.factset.com/hubfs/Website/Resources%20Section/Brochures/esg-data-and-analytics-from-truvalue-labs-brochure.pdf) Topic and category structure NOSIBLE Aspect-based sentiment is on the roadmap [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) FactSet 26 SASB categories [[ 2 a]](https://archive.is/PK6Ir) [[ 2 b]](https://web.archive.org/web/20260713190535/https://advantage.factset.com/hubfs/Website/Resources%20Section/Brochures/esg-data-and-analytics-from-truvalue-labs-brochure.pdf) Entity and ticker mapping NOSIBLE 3.2M organizations; 38,097 stable ticker mappings [[ 5 a]](https://archive.is/tGYc9) [[ 6 a]](https://archive.is/tnkpG) FactSet 260,000+ public and private companies [[ 1 a]](https://archive.is/eZ2mt) [[ 1 b]](https://web.archive.org/web/20221028182449/https://insight.factset.com/resources/at-a-glance-factset-truvalue-sasb-scores-datafeed) Ready-made signals NOSIBLE Raw dated events and search results [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) FactSet Pulse, Insight, Momentum, and Volume scores [[ 2 a]](https://archive.is/PK6Ir) [[ 2 b]](https://web.archive.org/web/20260713190535/https://advantage.factset.com/hubfs/Website/Resources%20Section/Brochures/esg-data-and-analytics-from-truvalue-labs-brochure.pdf) Delivery and API NOSIBLE Six-endpoint agent API, SDKs, MCP [[ 4 a]](https://archive.is/glqrw) FactSet FactSet data feed with daily delivery [[ 1 a]](https://archive.is/eZ2mt) [[ 1 b]](https://web.archive.org/web/20221028182449/https://insight.factset.com/resources/at-a-glance-factset-truvalue-sasb-scores-datafeed) Published research NOSIBLE Product materials describe event research and backtesting workflows [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) FactSet FactSet publishes a 2007–2019 score study [[ 1 a]](https://archive.is/eZ2mt) [[ 1 b]](https://web.archive.org/web/20221028182449/https://insight.factset.com/resources/at-a-glance-factset-truvalue-sasb-scores-datafeed) Pricing NOSIBLE See nosible.com [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) FactSet Contact FactSet for access and commercial terms [[ 1 a]](https://archive.is/eZ2mt) [[ 1 b]](https://web.archive.org/web/20221028182449/https://insight.factset.com/resources/at-a-glance-factset-truvalue-sasb-scores-datafeed) | Dimension | NOSIBLE | FactSet | | --- | --- | --- | | Category | Web-scale search and market-event intelligence for AI agents [[ 4 a]](https://archive.is/glqrw) | Outside-in ESG signals and scores [[ 1 a]](https://archive.is/eZ2mt) [[ 1 b]](https://web.archive.org/web/20221028182449/https://insight.factset.com/resources/at-a-glance-factset-truvalue-sasb-scores-datafeed) | | Content scope | Long-form news, corporate, government; excludes social, paywalled, marketplaces, adult [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | News, trade journals, NGO and industry reports (third-party, not self-reported) [[ 1 a]](https://archive.is/eZ2mt) [[ 1 b]](https://web.archive.org/web/20221028182449/https://insight.factset.com/resources/at-a-glance-factset-truvalue-sasb-scores-datafeed) | | Languages | 95 [[ 4 a]](https://archive.is/glqrw) | 30+ [[ 1 a]](https://archive.is/eZ2mt) [[ 1 b]](https://web.archive.org/web/20221028182449/https://insight.factset.com/resources/at-a-glance-factset-truvalue-sasb-scores-datafeed) | | History and point-in-time | About 30 years; five-way date-verification method [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Data available back to 2007; event-driven data delivered daily [[ 1 a]](https://archive.is/eZ2mt) [[ 1 b]](https://web.archive.org/web/20221028182449/https://insight.factset.com/resources/at-a-glance-factset-truvalue-sasb-scores-datafeed) | | Event database | Standalone 100M+ dated, ranked events (WORLD) [[ 5 a]](https://archive.is/tGYc9) | Positive and negative ESG events categorized by SASB topics [[ 2 a]](https://archive.is/PK6Ir) [[ 2 b]](https://web.archive.org/web/20260713190535/https://advantage.factset.com/hubfs/Website/Resources%20Section/Brochures/esg-data-and-analytics-from-truvalue-labs-brochure.pdf) | | Sentiment | Available in enrichment [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Continuously updated material ESG scores [[ 2 a]](https://archive.is/PK6Ir) [[ 2 b]](https://web.archive.org/web/20260713190535/https://advantage.factset.com/hubfs/Website/Resources%20Section/Brochures/esg-data-and-analytics-from-truvalue-labs-brochure.pdf) | | Topic and category structure | Aspect-based sentiment is on the roadmap [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | 26 SASB categories [[ 2 a]](https://archive.is/PK6Ir) [[ 2 b]](https://web.archive.org/web/20260713190535/https://advantage.factset.com/hubfs/Website/Resources%20Section/Brochures/esg-data-and-analytics-from-truvalue-labs-brochure.pdf) | | Entity and ticker mapping | 3.2M organizations; 38,097 stable ticker mappings [[ 5 a]](https://archive.is/tGYc9) [[ 6 a]](https://archive.is/tnkpG) | 260,000+ public and private companies [[ 1 a]](https://archive.is/eZ2mt) [[ 1 b]](https://web.archive.org/web/20221028182449/https://insight.factset.com/resources/at-a-glance-factset-truvalue-sasb-scores-datafeed) | | Ready-made signals | Raw dated events and search results [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Pulse, Insight, Momentum, and Volume scores [[ 2 a]](https://archive.is/PK6Ir) [[ 2 b]](https://web.archive.org/web/20260713190535/https://advantage.factset.com/hubfs/Website/Resources%20Section/Brochures/esg-data-and-analytics-from-truvalue-labs-brochure.pdf) | | Delivery and API | Six-endpoint agent API, SDKs, MCP [[ 4 a]](https://archive.is/glqrw) | FactSet data feed with daily delivery [[ 1 a]](https://archive.is/eZ2mt) [[ 1 b]](https://web.archive.org/web/20221028182449/https://insight.factset.com/resources/at-a-glance-factset-truvalue-sasb-scores-datafeed) | | Published research | Product materials describe event research and backtesting workflows [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | FactSet publishes a 2007–2019 score study [[ 1 a]](https://archive.is/eZ2mt) [[ 1 b]](https://web.archive.org/web/20221028182449/https://insight.factset.com/resources/at-a-glance-factset-truvalue-sasb-scores-datafeed) | | Pricing | See nosible.com [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Contact FactSet for access and commercial terms [[ 1 a]](https://archive.is/eZ2mt) [[ 1 b]](https://web.archive.org/web/20221028182449/https://insight.factset.com/resources/at-a-glance-factset-truvalue-sasb-scores-datafeed) | ## Potential fit by workflow Truvalue's cited materials describe SASB-aligned company ESG scores delivered through FactSet. [[ 2 a]](https://archive.is/PK6Ir) [[ 2 b]](https://web.archive.org/web/20260713190535/https://advantage.factset.com/hubfs/Website/Resources%20Section/Brochures/esg-data-and-analytics-from-truvalue-labs-brochure.pdf) NOSIBLE's cited materials describe cross-domain dated documents, ranked events, and agent-driven research. [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) In NOSIBLE's view, consider the product whose published output matches the requirement; the products may also be complementary when a finished ESG signal needs additional source context. ## Common FactSet Truvalue Labs comparison questions ### How does NOSIBLE feel it differentiates itself from FactSet? NOSIBLE is an AI-native company with two products: SEARCH and WORLD. [[ 5 a]](https://archive.is/tGYc9) [[ 7 a]](https://archive.is/iDZeo) SEARCH lets agents find dated open-web sources they can cite and inspect directly. [[ 4 a]](https://archive.is/glqrw) WORLD is a live open-web event database for models and backtests, with an embedding per event. [[ 5 a]](https://archive.is/tGYc9) [[ 6 a]](https://archive.is/tnkpG) [[ 8 a]](https://archive.is/1CUvQ) NOSIBLE is committed to open-source software and makes its models publicly available on Hugging Face. [[ 9 a]](https://archive.is/7SSz0) [[ 10 a]](https://archive.is/kHxMG) Related [WORLD v1.2 trial](https://nosible.com/start-trial#data-coverage) [Sentiment model](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) [Forward-looking model](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) ### If we need SASB-aligned ESG scores, can NOSIBLE replace FactSet Truvalue? Not as a drop-in replacement. [[ 1 a]](https://archive.is/eZ2mt) [[ 1 b]](https://web.archive.org/web/20221028182449/https://insight.factset.com/resources/at-a-glance-factset-truvalue-sasb-scores-datafeed) [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Truvalue supplies defined ESG scores and SASB categorization. [[ 2 a]](https://archive.is/PK6Ir) [[ 2 b]](https://web.archive.org/web/20260713190535/https://advantage.factset.com/hubfs/Website/Resources%20Section/Brochures/esg-data-and-analytics-from-truvalue-labs-brochure.pdf) NOSIBLE provides dated event evidence, ticker and entity mappings, SDG and risk tags, and enrichment tools from which a team can build a custom methodology. [[ 5 a]](https://archive.is/tGYc9) Related [WORLD event database](https://nosible.world/world) [Semantic factors](https://nosible.com/semantic-factors) ### Can NOSIBLE reproduce Truvalue's Pulse, Insight, Momentum, and Volume scores? No. Those are FactSet Truvalue products with their own methodology. [[ 1 a]](https://archive.is/eZ2mt) [[ 1 b]](https://web.archive.org/web/20221028182449/https://insight.factset.com/resources/at-a-glance-factset-truvalue-sasb-scores-datafeed) NOSIBLE can provide source-attributed events and features for separately designed indicators, but buyers should not describe those outputs as Truvalue scores or expect numerical equivalence. [[ 1 a]](https://archive.is/eZ2mt) [[ 1 b]](https://web.archive.org/web/20221028182449/https://insight.factset.com/resources/at-a-glance-factset-truvalue-sasb-scores-datafeed) [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Related [WORLD event database](https://nosible.world/world) ### How do Truvalue's outside-in ESG sources differ from NOSIBLE's corpus? Truvalue is ESG-first and includes external reporting, NGOs, watchdogs, trade publications, and social media. [[ 1 a]](https://archive.is/eZ2mt) [[ 1 b]](https://web.archive.org/web/20221028182449/https://insight.factset.com/resources/at-a-glance-factset-truvalue-sasb-scores-datafeed) NOSIBLE uses its own open-web source universe across ESG and non-ESG topics. [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Buyers should compare source overlap, licensing, document access, and historical availability directly. Related [WORLD event database](https://nosible.world/world) [Semantic factors](https://nosible.com/semantic-factors) ### What changes if our team already lives in FactSet Workstation and FactSet feeds? FactSet's cited materials describe delivery through its data-feed environment. [[ 1 a]](https://archive.is/eZ2mt) [[ 1 b]](https://web.archive.org/web/20221028182449/https://insight.factset.com/resources/at-a-glance-factset-truvalue-sasb-scores-datafeed) In NOSIBLE's view, that delivery may be the simpler option when existing systems already use FactSet identifiers and feeds. NOSIBLE adds an agent-oriented source-and-event workflow. [[ 4 a]](https://archive.is/glqrw) The integration decision should account for identifier mapping, entitlements, latency, and whether users need the underlying evidence. Related [WORLD event database](https://nosible.world/world) ### How does NOSIBLE handle ESG alpha and cross-domain event research? NOSIBLE treats ESG as one part of a cross-domain event model. [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Teams can combine sustainability evidence with policy, litigation, supply-chain, labor, reputational, and operational events. [[ 1 a]](https://archive.is/eZ2mt) [[ 1 b]](https://web.archive.org/web/20221028182449/https://insight.factset.com/resources/at-a-glance-factset-truvalue-sasb-scores-datafeed) [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Any claim of investment performance should be established through the buyer's own documented backtest rather than inferred from product coverage. Related [WORLD event database](https://nosible.world/world) [Semantic factors](https://nosible.com/semantic-factors) [Stock Screening](https://nosible.com/#proof) Continue comparing ## Related Vendor Comparisons Compare FactSet with adjacent options for task execution, agent tooling, data access, and research workflows, then evaluate the output shape your application actually needs for production. [RepRisk](https://nosible.com/compare/nosible-vs-reprisk) [SESAMm](https://nosible.com/compare/nosible-vs-sesamm) [Signal AI](https://nosible.com/compare/nosible-vs-signal-ai) [MarketPsych](https://nosible.com/compare/nosible-vs-marketpsych) Dive deeper ## Take your next step today Review the delivered field definitions, classification boundaries, and example values before comparing vendor workflows, so your team can assess what each product actually returns to downstream systems. [Data dictionaries](https://nosible.com/data-dictionaries) [Ontology reference](https://nosible.com/ontologies) [API reference](https://docs.nosible.com/) [Start trial](https://nosible.com/start-trial) Sources reviewed July 14, 2026 : [[ 1 ] FactSet Truvalue SASB Scores DataFeed](https://insight.factset.com/resources/at-a-glance-factset-truvalue-sasb-scores-datafeed) ( [dated snapshot](https://archive.is/eZ2mt); [Wayback copy](https://web.archive.org/web/20221028182449/https://insight.factset.com/resources/at-a-glance-factset-truvalue-sasb-scores-datafeed) ) , [[ 2 ] FactSet Truvalue ESG data brochure](https://advantage.factset.com/hubfs/Website/Resources%20Section/Brochures/esg-data-and-analytics-from-truvalue-labs-brochure.pdf) ( [dated snapshot](https://archive.is/PK6Ir); [Wayback copy](https://web.archive.org/web/20260713190535/https://advantage.factset.com/hubfs/Website/Resources%20Section/Brochures/esg-data-and-analytics-from-truvalue-labs-brochure.pdf) ) , [[ 3 ] NOSIBLE product overview](https://nosible.com/) ( [dated snapshot](https://archive.is/r7qDv); [Wayback copy](https://web.archive.org/web/20260713185114/https://nosible.com/) ) , [[ 4 ] NOSIBLE SEARCH](https://docs.nosible.com/) ( [dated snapshot](https://archive.is/glqrw) ) , [[ 5 ] NOSIBLE WORLD](https://nosible.world/world) ( [dated snapshot](https://archive.is/tGYc9) ) , [[ 6 ] NOSIBLE WORLD v1.2 trial and coverage](https://nosible.com/start-trial) ( [dated snapshot](https://archive.is/tnkpG) ) , [[ 7 ] NOSIBLE AI-native research overview](https://nosible.com/blog) ( [dated snapshot](https://archive.is/iDZeo) ) , [[ 8 ] NOSIBLE embedding-based research](https://nosible.com/blog/an-embedding-based-approach-to-trade-and-economic-policy-uncertainty) ( [dated snapshot](https://archive.is/1CUvQ) ) , [[ 9 ] NOSIBLE Financial Sentiment v1.2 Base](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) ( [dated snapshot](https://archive.is/7SSz0) ) , [[ 10 ] NOSIBLE Forward-Looking v1.2 Base](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) ( [dated snapshot](https://archive.is/kHxMG) ) . This page compares NOSIBLE WORLD with FactSet Truvalue Labs. FactSet, Truvalue, and related product names are marks of their respective owners. NOSIBLE is not affiliated with, sponsored by, or endorsed by FactSet. Comparison standards, legal context & corrections This comparison was prepared by NOSIBLE, which has a commercial interest in the products being compared. It is based on the cited public materials as they appeared on July 14, 2026. NOSIBLE has not tested every competitor feature, and the page is not a complete statement of either product. Products change; confirm current requirements, availability, and commercial terms with each vendor. Factual claims are attributed to cited first-party materials; evaluative statements reflect NOSIBLE's opinion. Each source includes a dated archive.is snapshot and, where the Internet Archive captured that URL, a timestamp-specific Wayback copy. Archive availability is controlled by those services. Competitor names and marks are used only to identify the products being compared. No affiliation, sponsorship, or endorsement is implied. The FTC says truthful, non-deceptive comparative advertising may identify competitors ( [policy](https://www.ftc.gov/legal-library/browse/statement-policy-regarding-comparative-advertising); [dated snapshot](https://archive.is/nZXY4); [Wayback copy](https://web.archive.org/web/20260215042003/https://www.ftc.gov/legal-library/browse/statement-policy-regarding-comparative-advertising) ). For one U.S. example of nominative-use analysis, see *New Kids on the Block v. News America Publishing, Inc.*, 971 F.2d 302, 308 (9th Cir. 1992) ( [opinion](https://law.justia.com/cases/federal/appellate-courts/F2/971/302/72076/); [dated snapshot](https://archive.is/dnk8j); [Wayback copy](https://web.archive.org/web/20250211213616/https://law.justia.com/cases/federal/appellate-courts/F2/971/302/72076/) ). If you represent FactSet and believe a factual statement is inaccurate, email [stuart@nosible.com](mailto:stuart@nosible.com?subject=Correction%20request%3A%20NOSIBLE%20WORLD%20vs%20FactSet) with the specific claim and a supporting first-party URL. NOSIBLE will review and correct substantiated errors. > How NOSIBLE WORLD, a web-scale search and market-event intelligence engine for AI agents, compares with FactSet Truvalue Labs, an outside-in ESG signals provider. Coverage, point-in-time integrity, scope, and use cases. **URL:** https://nosible.com/compare/nosible-vs-factset-truvalue --- --- title: "NOSIBLE WORLD vs RepRisk" description: "How NOSIBLE WORLD, a web-scale search and market-event intelligence engine for AI agents, compares with RepRisk, an ESG and business-conduct risk data provider. Coverage, point-in-time integrity, scope, and use cases." url: "https://nosible.com/compare/nosible-vs-reprisk" --- Comparison / Reviewed July 14, 2026 # NOSIBLE WORLD vs RepRisk RepRisk combines AI with analyst review to produce ESG and business-conduct risk data. [[ 2 a]](https://archive.is/nqZ8Z) [[ 2 b]](https://web.archive.org/web/20260513115412/https://www.reprisk.com/insights/resources/methodology) NOSIBLE WORLD is a domain-agnostic dated-document and event engine for agents, research, and custom risk models. [[ 1 a]](https://archive.is/XjYy1) [[ 1 b]](https://web.archive.org/web/20260527065032/https://www.reprisk.com/technology) [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) NOSIBLE-AUTHORED COMPARISON · NOSIBLE HAS A COMMERCIAL INTEREST IN THIS COMPARISON · FACTS ATTRIBUTED TO FIRST-PARTY VENDOR MATERIALS · EVALUATIVE STATEMENTS ARE NOSIBLE'S OPINION · REVIEWED JULY 14, 2026 · DUAL PUBLIC ARCHIVES WHERE SUPPORTED · [STANDARDS & CORRECTIONS](https://nosible.com/compare/nosible-vs-reprisk#comparison-standards). IF YOU REPRESENT REPRISK AND BELIEVE A FACTUAL STATEMENT IS INACCURATE, EMAIL [STUART@NOSIBLE.COM](mailto:stuart@nosible.com?subject=Correction%20request%3A%20NOSIBLE%20WORLD%20vs%20RepRisk) WITH THE SPECIFIC CLAIM AND A SUPPORTING FIRST-PARTY URL. NOSIBLE WILL REVIEW AND CORRECT SUBSTANTIATED ERRORS. - RepRisk says it screens 175,000+ public sources and stakeholders in 80 languages every day. [[ 2 a]](https://archive.is/nqZ8Z) [[ 2 b]](https://web.archive.org/web/20260513115412/https://www.reprisk.com/insights/resources/methodology) - Its process combines machine learning with review by 150+ analysts and senior-analyst quality assurance. [[ 2 a]](https://archive.is/nqZ8Z) [[ 2 b]](https://web.archive.org/web/20260513115412/https://www.reprisk.com/insights/resources/methodology) - RepRisk reports data from January 2007, 300,000+ risk assessments, and 108 factors mapped to UNGC, SASB, and the SDGs. [[ 2 a]](https://archive.is/nqZ8Z) [[ 2 b]](https://web.archive.org/web/20260513115412/https://www.reprisk.com/insights/resources/methodology) - RepRisk's methodology says it does not verify or validate reported allegations; its analysts check source and classification quality. [[ 2 a]](https://archive.is/nqZ8Z) [[ 2 b]](https://web.archive.org/web/20260513115412/https://www.reprisk.com/insights/resources/methodology) - RepRisk reports 348,000+ companies plus infrastructure projects; NOSIBLE WORLD v1.2 reports about 3.2 million organizations and 38,097 ticker mappings. [[ 5 a]](https://archive.is/tGYc9) [[ 6 a]](https://archive.is/tnkpG) These are not equivalent risk assessments. [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) ## Conduct-risk data and cross-domain event research RepRisk describes a specialized research process that combines large-scale screening, human review, quality assurance, and defined risk metrics. [[ 2 a]](https://archive.is/nqZ8Z) [[ 2 b]](https://web.archive.org/web/20260513115412/https://www.reprisk.com/insights/resources/methodology) NOSIBLE's cited materials describe dated market, policy, operational, geopolitical, sustainability, reputational, and company-event evidence for agents and custom research. [[ 1 a]](https://archive.is/XjYy1) [[ 1 b]](https://web.archive.org/web/20260527065032/https://www.reprisk.com/technology) [[ 4 a]](https://archive.is/glqrw) In NOSIBLE's view, NOSIBLE is the stronger fit when those cross-domain event categories are central to the workflow. ## Feature comparison The competitor column summarizes what the named vendor's cited first-party materials describe; the NOSIBLE column summarizes NOSIBLE's cited materials. Category NOSIBLE Web-scale search and market-event intelligence for AI agents [[ 4 a]](https://archive.is/glqrw) RepRisk ESG and business-conduct risk data [[ 1 a]](https://archive.is/XjYy1) [[ 1 b]](https://web.archive.org/web/20260527065032/https://www.reprisk.com/technology) Content scope NOSIBLE Long-form news, corporate, government; excludes social, paywalled, marketplaces, adult [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) RepRisk Print and online media, social media, blogs, government bodies, regulators, think tanks, newsletters, and other public sources [[ 1 a]](https://archive.is/XjYy1) [[ 1 b]](https://web.archive.org/web/20260527065032/https://www.reprisk.com/technology) Focus NOSIBLE Market, policy, operational, geopolitical, sustainability, reputational, and company events [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) RepRisk ESG, reputational, and business-conduct risk incidents [[ 1 a]](https://archive.is/XjYy1) [[ 1 b]](https://web.archive.org/web/20260527065032/https://www.reprisk.com/technology) Languages NOSIBLE 95 [[ 4 a]](https://archive.is/glqrw) RepRisk 80+ [[ 1 a]](https://archive.is/XjYy1) [[ 1 b]](https://web.archive.org/web/20260527065032/https://www.reprisk.com/technology) History and point-in-time NOSIBLE About 30 years; five-way date-verification method [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) RepRisk Data history from January 2007 with daily updates [[ 1 a]](https://archive.is/XjYy1) [[ 1 b]](https://web.archive.org/web/20260527065032/https://www.reprisk.com/technology) Event database NOSIBLE Standalone 100M+ dated, ranked events (WORLD) [[ 5 a]](https://archive.is/tGYc9) RepRisk 300,000+ risk assessments [[ 1 a]](https://archive.is/XjYy1) [[ 1 b]](https://web.archive.org/web/20260527065032/https://www.reprisk.com/technology) Sentiment NOSIBLE Available in enrichment [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) RepRisk Incident severity, source reach, and novelty [[ 1 a]](https://archive.is/XjYy1) [[ 1 b]](https://web.archive.org/web/20260527065032/https://www.reprisk.com/technology) Topic and risk taxonomy NOSIBLE Event, sustainability, and risk tags [[ 5 a]](https://archive.is/tGYc9) RepRisk 108 risk factors mapped to UNGC, SASB, and the SDGs [[ 2 a]](https://archive.is/nqZ8Z) [[ 2 b]](https://web.archive.org/web/20260513115412/https://www.reprisk.com/insights/resources/methodology) Entity and ticker mapping NOSIBLE 3.2M organizations; 38,097 stable ticker mappings [[ 5 a]](https://archive.is/tGYc9) [[ 6 a]](https://archive.is/tnkpG) RepRisk 348,000+ public and private companies plus infrastructure projects [[ 1 a]](https://archive.is/XjYy1) [[ 1 b]](https://web.archive.org/web/20260527065032/https://www.reprisk.com/technology) Processing approach NOSIBLE Automated event extraction and enrichment [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) RepRisk 150+ analysts review results; senior analysts perform quality assurance [[ 2 a]](https://archive.is/nqZ8Z) [[ 2 b]](https://web.archive.org/web/20260513115412/https://www.reprisk.com/insights/resources/methodology) Delivery and API NOSIBLE Six-endpoint agent API, SDKs, MCP [[ 4 a]](https://archive.is/glqrw) RepRisk Daily research and risk metrics [[ 1 a]](https://archive.is/XjYy1) [[ 1 b]](https://web.archive.org/web/20260527065032/https://www.reprisk.com/technology) Agent-readiness NOSIBLE Built for AI agents [[ 4 a]](https://archive.is/glqrw) RepRisk Curated dataset and proprietary risk metrics [[ 1 a]](https://archive.is/XjYy1) [[ 1 b]](https://web.archive.org/web/20260527065032/https://www.reprisk.com/technology) Distribution NOSIBLE Direct API and SDKs [[ 4 a]](https://archive.is/glqrw) RepRisk Standard and customized RepRisk metrics [[ 1 a]](https://archive.is/XjYy1) [[ 1 b]](https://web.archive.org/web/20260527065032/https://www.reprisk.com/technology) Pricing NOSIBLE See nosible.com [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) RepRisk Contact RepRisk for access and commercial terms [[ 1 a]](https://archive.is/XjYy1) [[ 1 b]](https://web.archive.org/web/20260527065032/https://www.reprisk.com/technology) | Dimension | NOSIBLE | RepRisk | | --- | --- | --- | | Category | Web-scale search and market-event intelligence for AI agents [[ 4 a]](https://archive.is/glqrw) | ESG and business-conduct risk data [[ 1 a]](https://archive.is/XjYy1) [[ 1 b]](https://web.archive.org/web/20260527065032/https://www.reprisk.com/technology) | | Content scope | Long-form news, corporate, government; excludes social, paywalled, marketplaces, adult [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Print and online media, social media, blogs, government bodies, regulators, think tanks, newsletters, and other public sources [[ 1 a]](https://archive.is/XjYy1) [[ 1 b]](https://web.archive.org/web/20260527065032/https://www.reprisk.com/technology) | | Focus | Market, policy, operational, geopolitical, sustainability, reputational, and company events [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | ESG, reputational, and business-conduct risk incidents [[ 1 a]](https://archive.is/XjYy1) [[ 1 b]](https://web.archive.org/web/20260527065032/https://www.reprisk.com/technology) | | Languages | 95 [[ 4 a]](https://archive.is/glqrw) | 80+ [[ 1 a]](https://archive.is/XjYy1) [[ 1 b]](https://web.archive.org/web/20260527065032/https://www.reprisk.com/technology) | | History and point-in-time | About 30 years; five-way date-verification method [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Data history from January 2007 with daily updates [[ 1 a]](https://archive.is/XjYy1) [[ 1 b]](https://web.archive.org/web/20260527065032/https://www.reprisk.com/technology) | | Event database | Standalone 100M+ dated, ranked events (WORLD) [[ 5 a]](https://archive.is/tGYc9) | 300,000+ risk assessments [[ 1 a]](https://archive.is/XjYy1) [[ 1 b]](https://web.archive.org/web/20260527065032/https://www.reprisk.com/technology) | | Sentiment | Available in enrichment [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Incident severity, source reach, and novelty [[ 1 a]](https://archive.is/XjYy1) [[ 1 b]](https://web.archive.org/web/20260527065032/https://www.reprisk.com/technology) | | Topic and risk taxonomy | Event, sustainability, and risk tags [[ 5 a]](https://archive.is/tGYc9) | 108 risk factors mapped to UNGC, SASB, and the SDGs [[ 2 a]](https://archive.is/nqZ8Z) [[ 2 b]](https://web.archive.org/web/20260513115412/https://www.reprisk.com/insights/resources/methodology) | | Entity and ticker mapping | 3.2M organizations; 38,097 stable ticker mappings [[ 5 a]](https://archive.is/tGYc9) [[ 6 a]](https://archive.is/tnkpG) | 348,000+ public and private companies plus infrastructure projects [[ 1 a]](https://archive.is/XjYy1) [[ 1 b]](https://web.archive.org/web/20260527065032/https://www.reprisk.com/technology) | | Processing approach | Automated event extraction and enrichment [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | 150+ analysts review results; senior analysts perform quality assurance [[ 2 a]](https://archive.is/nqZ8Z) [[ 2 b]](https://web.archive.org/web/20260513115412/https://www.reprisk.com/insights/resources/methodology) | | Delivery and API | Six-endpoint agent API, SDKs, MCP [[ 4 a]](https://archive.is/glqrw) | Daily research and risk metrics [[ 1 a]](https://archive.is/XjYy1) [[ 1 b]](https://web.archive.org/web/20260527065032/https://www.reprisk.com/technology) | | Agent-readiness | Built for AI agents [[ 4 a]](https://archive.is/glqrw) | Curated dataset and proprietary risk metrics [[ 1 a]](https://archive.is/XjYy1) [[ 1 b]](https://web.archive.org/web/20260527065032/https://www.reprisk.com/technology) | | Distribution | Direct API and SDKs [[ 4 a]](https://archive.is/glqrw) | Standard and customized RepRisk metrics [[ 1 a]](https://archive.is/XjYy1) [[ 1 b]](https://web.archive.org/web/20260527065032/https://www.reprisk.com/technology) | | Pricing | See nosible.com [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Contact RepRisk for access and commercial terms [[ 1 a]](https://archive.is/XjYy1) [[ 1 b]](https://web.archive.org/web/20260527065032/https://www.reprisk.com/technology) | ## Potential fit by workflow RepRisk's cited materials describe curated ESG and business-conduct risk assessments with analyst oversight. [[ 2 a]](https://archive.is/nqZ8Z) [[ 2 b]](https://web.archive.org/web/20260513115412/https://www.reprisk.com/insights/resources/methodology) NOSIBLE's cited materials describe cross-domain dated source retrieval and ranked events for agents. [[ 1 a]](https://archive.is/XjYy1) [[ 1 b]](https://web.archive.org/web/20260527065032/https://www.reprisk.com/technology) [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) In NOSIBLE's view, consider the product whose published output matches the requirement; the products may also be complementary when a risk score needs additional source context. ## Common RepRisk comparison questions ### How does NOSIBLE feel it differentiates itself from RepRisk? NOSIBLE is an AI-native company with two products: SEARCH and WORLD. [[ 5 a]](https://archive.is/tGYc9) [[ 7 a]](https://archive.is/iDZeo) SEARCH lets agents find dated open-web sources they can cite and inspect directly. [[ 4 a]](https://archive.is/glqrw) WORLD is a live open-web event database for models and backtests, with an embedding per event. [[ 5 a]](https://archive.is/tGYc9) [[ 6 a]](https://archive.is/tnkpG) [[ 8 a]](https://archive.is/1CUvQ) NOSIBLE is committed to open-source software and makes its models publicly available on Hugging Face. [[ 9 a]](https://archive.is/7SSz0) [[ 10 a]](https://archive.is/kHxMG) Related [WORLD v1.2 trial](https://nosible.com/start-trial#data-coverage) [Sentiment model](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) [Forward-looking model](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) ### Can NOSIBLE replace RepRisk for supplier, private-market, or infrastructure-project due diligence? RepRisk publishes defined company and infrastructure-project risk assessments. [[ 1 a]](https://archive.is/XjYy1) [[ 1 b]](https://web.archive.org/web/20260527065032/https://www.reprisk.com/technology) NOSIBLE can add dated source trails, entity context, and cross-domain events for investigation. [[ 1 a]](https://archive.is/XjYy1) [[ 1 b]](https://web.archive.org/web/20260527065032/https://www.reprisk.com/technology) [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) In NOSIBLE's view, RepRisk is the more direct choice when its defined assessment is required. NOSIBLE should not be described as a drop-in replacement without testing the required risk framework and coverage. [[ 1 a]](https://archive.is/XjYy1) [[ 1 b]](https://web.archive.org/web/20260527065032/https://www.reprisk.com/technology) [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Related [WORLD event database](https://nosible.world/world) [Stock Screening](https://nosible.com/#proof) ### How do RepRisk's RRI and RRR differ from NOSIBLE's event database? RepRisk's RRI and RRR are proprietary metrics derived from its curated incident process. [[ 1 a]](https://archive.is/XjYy1) [[ 1 b]](https://web.archive.org/web/20260527065032/https://www.reprisk.com/technology) NOSIBLE provides source-attributed events and entity mappings for custom models. [[ 1 a]](https://archive.is/XjYy1) [[ 1 b]](https://web.archive.org/web/20260527065032/https://www.reprisk.com/technology) [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) The outputs have different definitions and should not be treated as equivalent scores. [[ 1 a]](https://archive.is/XjYy1) [[ 1 b]](https://web.archive.org/web/20260527065032/https://www.reprisk.com/technology) Related [WORLD event database](https://nosible.world/world) ### Does NOSIBLE cover events RepRisk intentionally excludes? NOSIBLE's published event taxonomy includes positive, neutral, and adverse events across market, policy, macro, supply-chain, geopolitical, sustainability, reputational, and company contexts. [[ 1 a]](https://archive.is/XjYy1) [[ 1 b]](https://web.archive.org/web/20260527065032/https://www.reprisk.com/technology) [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) In NOSIBLE's view, that cross-domain stated scope is useful for research, while RepRisk's specialized screening and curation remain distinct advantages for conduct-risk use cases. Related [WORLD event database](https://nosible.world/world) [Geopolitical risk index](https://nosible.com/blog/rebuilding-the-geopolitical-risk-index-from-nosible-world) [Stock Screening](https://nosible.com/#proof) ### How does NOSIBLE fit around UNGC, SASB, SFDR, Modern Slavery, or SDG mapping? RepRisk publishes 108 factors mapped to UNGC, SASB, and the SDGs. [[ 2 a]](https://archive.is/nqZ8Z) [[ 2 b]](https://web.archive.org/web/20260513115412/https://www.reprisk.com/insights/resources/methodology) NOSIBLE supplies dated evidence and its own sustainability and risk tags for custom models. [[ 1 a]](https://archive.is/XjYy1) [[ 1 b]](https://web.archive.org/web/20260527065032/https://www.reprisk.com/technology) [[ 5 a]](https://archive.is/tGYc9) Buyers needing a specific regulatory framework should validate exact field definitions and methodology with each vendor. [[ 1 a]](https://archive.is/XjYy1) [[ 1 b]](https://web.archive.org/web/20260527065032/https://www.reprisk.com/technology) Related [WORLD event database](https://nosible.world/world) [Semantic factors](https://nosible.com/semantic-factors) ### How should buyers verify or challenge an incident before acting on it? RepRisk states that analysts review incidents and senior analysts perform quality assurance. [[ 2 a]](https://archive.is/nqZ8Z) [[ 2 b]](https://web.archive.org/web/20260513115412/https://www.reprisk.com/insights/resources/methodology) Buyers should still inspect the available source trail, date, entity match, severity, reach, and novelty before acting. NOSIBLE can support additional source retrieval and contextual research around an incident. [[ 1 a]](https://archive.is/XjYy1) [[ 1 b]](https://web.archive.org/web/20260527065032/https://www.reprisk.com/technology) [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Related [WORLD event database](https://nosible.world/world) Continue comparing ## Related Vendor Comparisons Compare RepRisk with adjacent options for task execution, agent tooling, data access, and research workflows, then evaluate the output shape your application actually needs for production. [FactSet Truvalue](https://nosible.com/compare/nosible-vs-factset-truvalue) [SESAMm](https://nosible.com/compare/nosible-vs-sesamm) [Signal AI](https://nosible.com/compare/nosible-vs-signal-ai) [Bloomberg](https://nosible.com/compare/nosible-vs-bloomberg) Dive deeper ## Take your next step today Review the delivered field definitions, classification boundaries, and example values before comparing vendor workflows, so your team can assess what each product actually returns to downstream systems. [Data dictionaries](https://nosible.com/data-dictionaries) [Ontology reference](https://nosible.com/ontologies) [API reference](https://docs.nosible.com/) [Start trial](https://nosible.com/start-trial) Sources reviewed July 14, 2026 : [[ 1 ] RepRisk technology and methodology](https://www.reprisk.com/technology) ( [dated snapshot](https://archive.is/XjYy1); [Wayback copy](https://web.archive.org/web/20260527065032/https://www.reprisk.com/technology) ) , [[ 2 ] RepRisk methodology](https://www.reprisk.com/insights/resources/methodology) ( [dated snapshot](https://archive.is/nqZ8Z); [Wayback copy](https://web.archive.org/web/20260513115412/https://www.reprisk.com/insights/resources/methodology) ) , [[ 3 ] NOSIBLE product overview](https://nosible.com/) ( [dated snapshot](https://archive.is/r7qDv); [Wayback copy](https://web.archive.org/web/20260713185114/https://nosible.com/) ) , [[ 4 ] NOSIBLE SEARCH](https://docs.nosible.com/) ( [dated snapshot](https://archive.is/glqrw) ) , [[ 5 ] NOSIBLE WORLD](https://nosible.world/world) ( [dated snapshot](https://archive.is/tGYc9) ) , [[ 6 ] NOSIBLE WORLD v1.2 trial and coverage](https://nosible.com/start-trial) ( [dated snapshot](https://archive.is/tnkpG) ) , [[ 7 ] NOSIBLE AI-native research overview](https://nosible.com/blog) ( [dated snapshot](https://archive.is/iDZeo) ) , [[ 8 ] NOSIBLE embedding-based research](https://nosible.com/blog/an-embedding-based-approach-to-trade-and-economic-policy-uncertainty) ( [dated snapshot](https://archive.is/1CUvQ) ) , [[ 9 ] NOSIBLE Financial Sentiment v1.2 Base](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) ( [dated snapshot](https://archive.is/7SSz0) ) , [[ 10 ] NOSIBLE Forward-Looking v1.2 Base](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) ( [dated snapshot](https://archive.is/kHxMG) ) . This page compares NOSIBLE WORLD with RepRisk. RepRisk and related product names are marks of their respective owner. NOSIBLE is not affiliated with, sponsored by, or endorsed by RepRisk. Comparison standards, legal context & corrections This comparison was prepared by NOSIBLE, which has a commercial interest in the products being compared. It is based on the cited public materials as they appeared on July 14, 2026. NOSIBLE has not tested every competitor feature, and the page is not a complete statement of either product. Products change; confirm current requirements, availability, and commercial terms with each vendor. Factual claims are attributed to cited first-party materials; evaluative statements reflect NOSIBLE's opinion. Each source includes a dated archive.is snapshot and, where the Internet Archive captured that URL, a timestamp-specific Wayback copy. Archive availability is controlled by those services. Competitor names and marks are used only to identify the products being compared. No affiliation, sponsorship, or endorsement is implied. The FTC says truthful, non-deceptive comparative advertising may identify competitors ( [policy](https://www.ftc.gov/legal-library/browse/statement-policy-regarding-comparative-advertising); [dated snapshot](https://archive.is/nZXY4); [Wayback copy](https://web.archive.org/web/20260215042003/https://www.ftc.gov/legal-library/browse/statement-policy-regarding-comparative-advertising) ). For one U.S. example of nominative-use analysis, see *New Kids on the Block v. News America Publishing, Inc.*, 971 F.2d 302, 308 (9th Cir. 1992) ( [opinion](https://law.justia.com/cases/federal/appellate-courts/F2/971/302/72076/); [dated snapshot](https://archive.is/dnk8j); [Wayback copy](https://web.archive.org/web/20250211213616/https://law.justia.com/cases/federal/appellate-courts/F2/971/302/72076/) ). If you represent RepRisk and believe a factual statement is inaccurate, email [stuart@nosible.com](mailto:stuart@nosible.com?subject=Correction%20request%3A%20NOSIBLE%20WORLD%20vs%20RepRisk) with the specific claim and a supporting first-party URL. NOSIBLE will review and correct substantiated errors. > How NOSIBLE WORLD, a web-scale search and market-event intelligence engine for AI agents, compares with RepRisk, an ESG and business-conduct risk data provider. Coverage, point-in-time integrity, scope, and use cases. **URL:** https://nosible.com/compare/nosible-vs-reprisk --- --- title: "NOSIBLE vs SESAMm" description: "Compare NOSIBLE and SESAMm on point-in-time integrity, event intelligence, ESG workflows, source coverage, languages, APIs, and AI-agent access." url: "https://nosible.com/compare/nosible-vs-sesamm" --- Comparison / Reviewed July 14, 2026 # NOSIBLE vs SESAMm SESAMm analyzes web data for ESG, reputational, KYC, supplier, and other risk workflows. [[ 2 a]](https://archive.is/23ODA) [[ 2 b]](https://web.archive.org/web/20260418150140/https://www.sesamm.com/) NOSIBLE is a cross-domain dated source-and-event engine for agents and research. [[ 4 a]](https://archive.is/glqrw) NOSIBLE-AUTHORED COMPARISON · NOSIBLE HAS A COMMERCIAL INTEREST IN THIS COMPARISON · FACTS ATTRIBUTED TO FIRST-PARTY VENDOR MATERIALS · EVALUATIVE STATEMENTS ARE NOSIBLE'S OPINION · REVIEWED JULY 14, 2026 · DUAL PUBLIC ARCHIVES WHERE SUPPORTED · [STANDARDS & CORRECTIONS](https://nosible.com/compare/nosible-vs-sesamm#comparison-standards). IF YOU REPRESENT SESAMM AND BELIEVE A FACTUAL STATEMENT IS INACCURATE, EMAIL [STUART@NOSIBLE.COM](mailto:stuart@nosible.com?subject=Correction%20request%3A%20NOSIBLE%20vs%20SESAMm) WITH THE SPECIFIC CLAIM AND A SUPPORTING FIRST-PARTY URL. NOSIBLE WILL REVIEW AND CORRECT SUBSTANTIATED ERRORS. - SESAMm's TextReveal docs page says it processes 20+ billion historical articles and 10+ million new documents daily; its current homepage separately reports 30B+ documents overall. [[ 2 a]](https://archive.is/23ODA) [[ 2 b]](https://web.archive.org/web/20260418150140/https://www.sesamm.com/) - SESAMm reports 5M+ companies; NOSIBLE WORLD v1.2 reports about 3.2 million organizations and 6.4 million people. [[ 5 a]](https://archive.is/tGYc9) [[ 6 a]](https://archive.is/tnkpG) The vendors' published entity definitions may differ. [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) - TextReveal provides ESG and SDG events through alerts, dashboards, APIs, and flat files with links to source articles. [[ 1 a]](https://archive.is/Whjyk) - SESAMm's current homepage also describes AI-agent, KYC, supplier, legal, regulatory, and reputational-risk workflows. [[ 2 a]](https://archive.is/23ODA) [[ 2 b]](https://web.archive.org/web/20260418150140/https://www.sesamm.com/) - NOSIBLE emphasizes dated retrieval, ranked evidence, and cross-domain agent workflows. [[ 4 a]](https://archive.is/glqrw) - In NOSIBLE's view, consider NOSIBLE when market-event research rather than a defined risk workflow is the primary requirement. ## Defined risk workflows and cross-domain event research SESAMm publishes ESG, reputational, KYC, supplier, legal, regulatory, and related risk workflows with multilingual coverage and source-backed outputs. [[ 2 a]](https://archive.is/23ODA) [[ 2 b]](https://web.archive.org/web/20260418150140/https://www.sesamm.com/) NOSIBLE focuses on a market-intelligence layer: dated source material, ranked events, and agent-oriented retrieval across many event categories. [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) In NOSIBLE's view, NOSIBLE is the stronger fit when market-event research is the defining requirement. ## Risk workflows and cross-domain event research The competitor column summarizes what the named vendor's cited first-party materials describe; the NOSIBLE column summarizes NOSIBLE's cited materials. Primary use NOSIBLE Search and market-event intelligence for AI agents [[ 4 a]](https://archive.is/glqrw) SESAMm ESG, reputational, KYC, supplier, legal, and regulatory risk workflows [[ 2 a]](https://archive.is/23ODA) [[ 2 b]](https://web.archive.org/web/20260418150140/https://www.sesamm.com/) Illustrative users NOSIBLE Quant, risk, agents, event research [[ 4 a]](https://archive.is/glqrw) SESAMm ESG, reputational-risk, KYC, supplier-risk, compliance, and investment teams [[ 2 a]](https://archive.is/23ODA) [[ 2 b]](https://web.archive.org/web/20260418150140/https://www.sesamm.com/) Content scope NOSIBLE Long-form news, corporate, and government text [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) SESAMm Web documents analyzed for ESG, reputational, KYC, supplier, legal, regulatory, and related risk workflows [[ 2 a]](https://archive.is/23ODA) [[ 2 b]](https://web.archive.org/web/20260418150140/https://www.sesamm.com/) History NOSIBLE Roughly 30 years [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) SESAMm Cited pages: 20B+ historical articles; current homepage: 30B+ documents overall [[ 2 a]](https://archive.is/23ODA) [[ 2 b]](https://web.archive.org/web/20260418150140/https://www.sesamm.com/) Languages NOSIBLE 95 [[ 4 a]](https://archive.is/glqrw) SESAMm 100+ [[ 1 a]](https://archive.is/Whjyk) Point-in-time method NOSIBLE Five-way date verification [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) SESAMm Near-live events with access to original articles [[ 1 a]](https://archive.is/Whjyk) Event data NOSIBLE WORLD, 100M+ ranked dated events [[ 5 a]](https://archive.is/tGYc9) SESAMm ESG, reputational, legal, regulatory, supplier, and related risk events [[ 2 a]](https://archive.is/23ODA) [[ 2 b]](https://web.archive.org/web/20260418150140/https://www.sesamm.com/) Sentiment NOSIBLE Enrichment layer within cross-domain event intelligence [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) SESAMm Sentiment analysis and risk-event classification [[ 1 a]](https://archive.is/Whjyk) Delivery NOSIBLE Six-endpoint API, SDKs, MCP, Cybernaut-1 [[ 4 a]](https://archive.is/glqrw) SESAMm Email alerts, dashboards, API, and flat-file delivery [[ 1 a]](https://archive.is/Whjyk) Agent readiness NOSIBLE Built agent-first [[ 4 a]](https://archive.is/glqrw) SESAMm API documentation and MCP server [[ 1 a]](https://archive.is/Whjyk) | Dimension | NOSIBLE | SESAMm | | --- | --- | --- | | Primary use | Search and market-event intelligence for AI agents [[ 4 a]](https://archive.is/glqrw) | ESG, reputational, KYC, supplier, legal, and regulatory risk workflows [[ 2 a]](https://archive.is/23ODA) [[ 2 b]](https://web.archive.org/web/20260418150140/https://www.sesamm.com/) | | Illustrative users | Quant, risk, agents, event research [[ 4 a]](https://archive.is/glqrw) | ESG, reputational-risk, KYC, supplier-risk, compliance, and investment teams [[ 2 a]](https://archive.is/23ODA) [[ 2 b]](https://web.archive.org/web/20260418150140/https://www.sesamm.com/) | | Content scope | Long-form news, corporate, and government text [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Web documents analyzed for ESG, reputational, KYC, supplier, legal, regulatory, and related risk workflows [[ 2 a]](https://archive.is/23ODA) [[ 2 b]](https://web.archive.org/web/20260418150140/https://www.sesamm.com/) | | History | Roughly 30 years [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Cited pages: 20B+ historical articles; current homepage: 30B+ documents overall [[ 2 a]](https://archive.is/23ODA) [[ 2 b]](https://web.archive.org/web/20260418150140/https://www.sesamm.com/) | | Languages | 95 [[ 4 a]](https://archive.is/glqrw) | 100+ [[ 1 a]](https://archive.is/Whjyk) | | Point-in-time method | Five-way date verification [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Near-live events with access to original articles [[ 1 a]](https://archive.is/Whjyk) | | Event data | WORLD, 100M+ ranked dated events [[ 5 a]](https://archive.is/tGYc9) | ESG, reputational, legal, regulatory, supplier, and related risk events [[ 2 a]](https://archive.is/23ODA) [[ 2 b]](https://web.archive.org/web/20260418150140/https://www.sesamm.com/) | | Sentiment | Enrichment layer within cross-domain event intelligence [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Sentiment analysis and risk-event classification [[ 1 a]](https://archive.is/Whjyk) | | Delivery | Six-endpoint API, SDKs, MCP, Cybernaut-1 [[ 4 a]](https://archive.is/glqrw) | Email alerts, dashboards, API, and flat-file delivery [[ 1 a]](https://archive.is/Whjyk) | | Agent readiness | Built agent-first [[ 4 a]](https://archive.is/glqrw) | API documentation and MCP server [[ 1 a]](https://archive.is/Whjyk) | ## Potential fit by workflow SESAMm's cited materials describe ESG, reputational, KYC, supplier, and related risk products. [[ 2 a]](https://archive.is/23ODA) [[ 2 b]](https://web.archive.org/web/20260418150140/https://www.sesamm.com/) NOSIBLE's cited materials describe dated source discovery and ranked events across market, policy, operational, geopolitical, and sustainability topics. [[ 5 a]](https://archive.is/tGYc9) In NOSIBLE's view, consider the product whose published workflow matches the requirement. Source overlap and timing should be tested directly. ## Common SESAMm comparison questions ### How does NOSIBLE feel it differentiates itself from SESAMm? NOSIBLE is an AI-native company with two products: SEARCH and WORLD. [[ 5 a]](https://archive.is/tGYc9) [[ 7 a]](https://archive.is/iDZeo) SEARCH lets agents find dated open-web sources they can cite and inspect directly. [[ 4 a]](https://archive.is/glqrw) WORLD is a live open-web event database for models and backtests, with an embedding per event. [[ 5 a]](https://archive.is/tGYc9) [[ 6 a]](https://archive.is/tnkpG) [[ 8 a]](https://archive.is/1CUvQ) NOSIBLE is committed to open-source software and makes its models publicly available on Hugging Face. [[ 9 a]](https://archive.is/7SSz0) [[ 10 a]](https://archive.is/kHxMG) Related [WORLD v1.2 trial](https://nosible.com/start-trial#data-coverage) [Sentiment model](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) [Forward-looking model](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) ### Is NOSIBLE an ESG ratings provider like SESAMm? SESAMm publishes defined ESG and SDG monitoring products. [[ 1 a]](https://archive.is/Whjyk) NOSIBLE is not sold as a drop-in SESAMm score; it provides dated source evidence, ranked events, entity mappings, and sustainability and risk tags that teams can use in custom models. [[ 1 a]](https://archive.is/Whjyk) [[ 5 a]](https://archive.is/tGYc9) In NOSIBLE's view, SESAMm is the more direct choice when one of those defined products is required. Related [WORLD event database](https://nosible.world/world) [Semantic factors](https://nosible.com/semantic-factors) ### When should we choose NOSIBLE if both products have APIs and AI-agent workflows? SESAMm documents API and MCP access for its risk data. [[ 1 a]](https://archive.is/Whjyk) NOSIBLE's agent workflow is oriented toward cross-domain search, source history, and ranked events. [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) In NOSIBLE's view, choose based on whether the agent primarily consumes a defined ESG dataset or must investigate new questions across a cross-domain source universe. Related [Agentic Search](https://docs.nosible.com/) [WORLD event database](https://nosible.world/world) [Semantic factors](https://nosible.com/semantic-factors) ### How do SESAMm's and NOSIBLE's source models differ? The products publish different source models and neither is universally preferable. [[ 1 a]](https://archive.is/Whjyk) [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) SESAMm publishes large multilingual document, source, and company counts for risk monitoring. [[ 1 a]](https://archive.is/Whjyk) NOSIBLE emphasizes dated, inspectable long-form sources and cross-domain event research. [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Test recall, precision, source access, duplication, and revision handling on a representative portfolio. Related [WORLD event database](https://nosible.world/world) ### How does SESAMm's Controversy Exposure Score differ from NOSIBLE WORLD events? SESAMm offers a defined Controversy Exposure Score and ESG event framework. [[ 1 a]](https://archive.is/Whjyk) NOSIBLE provides a cross-domain ranked-event layer across policy, macro, company, supply-chain, geopolitical, operational, sustainability, and reputational topics. [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) The outputs are not numerically interchangeable. [[ 1 a]](https://archive.is/Whjyk) [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Related [WORLD event database](https://nosible.world/world) [Geopolitical risk index](https://nosible.com/blog/rebuilding-the-geopolitical-risk-index-from-nosible-world) [Stock Screening](https://nosible.com/#proof) ### Can SESAMm support alpha backtests the same way NOSIBLE is positioned to? SESAMm's public documentation describes near-live updates and historical articles, but product suitability for a specific backtest must be tested. [[ 1 a]](https://archive.is/Whjyk) Ask both vendors about timestamps, revisions, entity history, and reproducible as-of queries, then run the same historical protocol on each dataset. Related [WORLD event database](https://nosible.world/world) [Bulk Web Search](https://docs.nosible.com/endpoints/search-bulk) Continue comparing ## Related Vendor Comparisons Compare SESAMm with adjacent options for task execution, agent tooling, data access, and research workflows, then evaluate the output shape your application actually needs for production. [FactSet Truvalue](https://nosible.com/compare/nosible-vs-factset-truvalue) [RepRisk](https://nosible.com/compare/nosible-vs-reprisk) [Signal AI](https://nosible.com/compare/nosible-vs-signal-ai) [MKT MediaStats](https://nosible.com/compare/nosible-vs-mkt-mediastats) Dive deeper ## Take your next step today Review the delivered field definitions, classification boundaries, and example values before comparing vendor workflows, so your team can assess what each product actually returns to downstream systems. [Data dictionaries](https://nosible.com/data-dictionaries) [Ontology reference](https://nosible.com/ontologies) [API reference](https://docs.nosible.com/) [Start trial](https://nosible.com/start-trial) Sources reviewed July 14, 2026 : [[ 1 ] TextReveal overview](https://docs.textreveal.com/guide/overview) ( [dated snapshot](https://archive.is/Whjyk) ) , [[ 2 ] SESAMm product overview](https://www.sesamm.com/) ( [dated snapshot](https://archive.is/23ODA); [Wayback copy](https://web.archive.org/web/20260418150140/https://www.sesamm.com/) ) , [[ 3 ] NOSIBLE product overview](https://nosible.com/) ( [dated snapshot](https://archive.is/r7qDv); [Wayback copy](https://web.archive.org/web/20260713185114/https://nosible.com/) ) , [[ 4 ] NOSIBLE SEARCH](https://docs.nosible.com/) ( [dated snapshot](https://archive.is/glqrw) ) , [[ 5 ] NOSIBLE WORLD](https://nosible.world/world) ( [dated snapshot](https://archive.is/tGYc9) ) , [[ 6 ] NOSIBLE WORLD v1.2 trial and coverage](https://nosible.com/start-trial) ( [dated snapshot](https://archive.is/tnkpG) ) , [[ 7 ] NOSIBLE AI-native research overview](https://nosible.com/blog) ( [dated snapshot](https://archive.is/iDZeo) ) , [[ 8 ] NOSIBLE embedding-based research](https://nosible.com/blog/an-embedding-based-approach-to-trade-and-economic-policy-uncertainty) ( [dated snapshot](https://archive.is/1CUvQ) ) , [[ 9 ] NOSIBLE Financial Sentiment v1.2 Base](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) ( [dated snapshot](https://archive.is/7SSz0) ) , [[ 10 ] NOSIBLE Forward-Looking v1.2 Base](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) ( [dated snapshot](https://archive.is/kHxMG) ) . SESAMm and TextReveal are trademarks of their respective owners. NOSIBLE is not affiliated with, sponsored by, or endorsed by SESAMm. Comparison standards, legal context & corrections This comparison was prepared by NOSIBLE, which has a commercial interest in the products being compared. It is based on the cited public materials as they appeared on July 14, 2026. NOSIBLE has not tested every competitor feature, and the page is not a complete statement of either product. Products change; confirm current requirements, availability, and commercial terms with each vendor. Factual claims are attributed to cited first-party materials; evaluative statements reflect NOSIBLE's opinion. Each source includes a dated archive.is snapshot and, where the Internet Archive captured that URL, a timestamp-specific Wayback copy. Archive availability is controlled by those services. Competitor names and marks are used only to identify the products being compared. No affiliation, sponsorship, or endorsement is implied. The FTC says truthful, non-deceptive comparative advertising may identify competitors ( [policy](https://www.ftc.gov/legal-library/browse/statement-policy-regarding-comparative-advertising); [dated snapshot](https://archive.is/nZXY4); [Wayback copy](https://web.archive.org/web/20260215042003/https://www.ftc.gov/legal-library/browse/statement-policy-regarding-comparative-advertising) ). For one U.S. example of nominative-use analysis, see *New Kids on the Block v. News America Publishing, Inc.*, 971 F.2d 302, 308 (9th Cir. 1992) ( [opinion](https://law.justia.com/cases/federal/appellate-courts/F2/971/302/72076/); [dated snapshot](https://archive.is/dnk8j); [Wayback copy](https://web.archive.org/web/20250211213616/https://law.justia.com/cases/federal/appellate-courts/F2/971/302/72076/) ). If you represent SESAMm and believe a factual statement is inaccurate, email [stuart@nosible.com](mailto:stuart@nosible.com?subject=Correction%20request%3A%20NOSIBLE%20vs%20SESAMm) with the specific claim and a supporting first-party URL. NOSIBLE will review and correct substantiated errors. > Compare NOSIBLE and SESAMm on point-in-time integrity, event intelligence, ESG workflows, source coverage, languages, APIs, and AI-agent access. **URL:** https://nosible.com/compare/nosible-vs-sesamm --- --- title: "NOSIBLE WORLD vs Signal AI" description: "How NOSIBLE WORLD, a web-scale search and market-event intelligence engine for AI agents, compares with Signal AI, a reputation and risk intelligence platform. Coverage, point-in-time integrity, scope, and use cases." url: "https://nosible.com/compare/nosible-vs-signal-ai" --- Comparison / Reviewed July 14, 2026 # NOSIBLE WORLD vs Signal AI Signal AI provides reputation and risk intelligence through APIs, dashboards, reports, and an AI agent. [[ 2 a]](https://archive.is/YQkKk) [[ 2 b]](https://web.archive.org/web/20260702044745/https://signal-ai.com/about/) NOSIBLE WORLD focuses on dated open-web sources and ranked events for investment, risk, and research agents. [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) NOSIBLE-AUTHORED COMPARISON · NOSIBLE HAS A COMMERCIAL INTEREST IN THIS COMPARISON · FACTS ATTRIBUTED TO FIRST-PARTY VENDOR MATERIALS · EVALUATIVE STATEMENTS ARE NOSIBLE'S OPINION · REVIEWED JULY 14, 2026 · DUAL PUBLIC ARCHIVES WHERE SUPPORTED · [STANDARDS & CORRECTIONS](https://nosible.com/compare/nosible-vs-signal-ai#comparison-standards). IF YOU REPRESENT SIGNAL AI AND BELIEVE A FACTUAL STATEMENT IS INACCURATE, EMAIL [STUART@NOSIBLE.COM](mailto:stuart@nosible.com?subject=Correction%20request%3A%20NOSIBLE%20WORLD%20vs%20Signal%20AI) WITH THE SPECIFIC CLAIM AND A SUPPORTING FIRST-PARTY URL. NOSIBLE WILL REVIEW AND CORRECT SUBSTANTIATED ERRORS. - Signal AI's cited API page says it ingests 5 million+ articles daily and covers and translates 75+ languages. [[ 1 a]](https://archive.is/98DLH) [[ 1 b]](https://web.archive.org/web/20260517020355/https://signal-ai.com/solutions/api/) - Its four published endpoints are Search, Metrics, Affinity, and Events. [[ 1 a]](https://archive.is/98DLH) [[ 1 b]](https://web.archive.org/web/20260517020355/https://signal-ai.com/solutions/api/) - Signal AI's Search endpoint covers more than a billion documents, and its current About page says its data can be accessed through intelligent agents. [[ 1 a]](https://archive.is/98DLH) [[ 1 b]](https://web.archive.org/web/20260517020355/https://signal-ai.com/solutions/api/) - NOSIBLE WORLD emphasizes dated source retrieval, ranked events, ticker context, and research-agent workflows. [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) ## Reputation intelligence and event research Signal AI describes an enterprise-intelligence platform with search, metrics, affinity, and event APIs plus ready-made reputation and risk workflows. [[ 2 a]](https://archive.is/YQkKk) [[ 2 b]](https://web.archive.org/web/20260702044745/https://signal-ai.com/about/) NOSIBLE WORLD is designed around dated open-web sources, ranked events, and securities-oriented research agents. [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) In NOSIBLE's view, NOSIBLE is the stronger fit when historical source evidence and investment context are the defining requirements. ## Feature comparison The competitor column summarizes what the named vendor's cited first-party materials describe; the NOSIBLE column summarizes NOSIBLE's cited materials. Category NOSIBLE Web-scale search and market-event intelligence for AI agents [[ 4 a]](https://archive.is/glqrw) Signal AI Reputation and risk intelligence platform [[ 2 a]](https://archive.is/YQkKk) [[ 2 b]](https://web.archive.org/web/20260702044745/https://signal-ai.com/about/) Illustrative users NOSIBLE Investors, quants, AI agents [[ 4 a]](https://archive.is/glqrw) Signal AI Reputation, communications, ESG, supply-chain, and enterprise-risk teams [[ 2 a]](https://archive.is/YQkKk) [[ 2 b]](https://web.archive.org/web/20260702044745/https://signal-ai.com/about/) Content scope NOSIBLE Long-form news, corporate, government; excludes social, paywalled, marketplaces, adult [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Signal AI Global news data including premium licensed content [[ 1 a]](https://archive.is/98DLH) [[ 1 b]](https://web.archive.org/web/20260517020355/https://signal-ai.com/solutions/api/) Languages NOSIBLE 95 [[ 4 a]](https://archive.is/glqrw) Signal AI Cited API page: 75+ covered and translated languages [[ 1 a]](https://archive.is/98DLH) [[ 1 b]](https://web.archive.org/web/20260517020355/https://signal-ai.com/solutions/api/) Published corpus indicators NOSIBLE 300,000+ open-web sources with source history [[ 4 a]](https://archive.is/glqrw) Signal AI Search across more than one billion documents [[ 1 a]](https://archive.is/98DLH) [[ 1 b]](https://web.archive.org/web/20260517020355/https://signal-ai.com/solutions/api/) Event database NOSIBLE Standalone 100M+ dated, ranked events (WORLD) [[ 5 a]](https://archive.is/tGYc9) Signal AI Events endpoint for monitoring and historical deep dives [[ 1 a]](https://archive.is/98DLH) [[ 1 b]](https://web.archive.org/web/20260517020355/https://signal-ai.com/solutions/api/) Analytical measures NOSIBLE Available in enrichment [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Signal AI Metrics endpoint for trends, patterns, and themes [[ 1 a]](https://archive.is/98DLH) [[ 1 b]](https://web.archive.org/web/20260517020355/https://signal-ai.com/solutions/api/) Relationship and aspect analysis NOSIBLE Aspect-based sentiment is on the roadmap [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Signal AI Affinity endpoint for relationships between concepts [[ 1 a]](https://archive.is/98DLH) [[ 1 b]](https://web.archive.org/web/20260517020355/https://signal-ai.com/solutions/api/) Entity and ticker mapping NOSIBLE Source-attributed entities [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Signal AI Topic and concept metadata for reputation analysis [[ 2 a]](https://archive.is/YQkKk) [[ 2 b]](https://web.archive.org/web/20260702044745/https://signal-ai.com/about/) Delivery and API NOSIBLE Six-endpoint agent API, SDKs, MCP [[ 4 a]](https://archive.is/glqrw) Signal AI API access for BI solutions and AI agents [[ 2 a]](https://archive.is/YQkKk) [[ 2 b]](https://web.archive.org/web/20260702044745/https://signal-ai.com/about/) Agent-readiness NOSIBLE Built for AI agents [[ 4 a]](https://archive.is/glqrw) Signal AI Direct support for AI-agent integration [[ 1 a]](https://archive.is/98DLH) [[ 1 b]](https://web.archive.org/web/20260517020355/https://signal-ai.com/solutions/api/) Published use cases NOSIBLE Product materials describe event research and backtesting workflows [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Signal AI Published use cases include reputation, ESG perception, and supply-chain risk [[ 2 a]](https://archive.is/YQkKk) [[ 2 b]](https://web.archive.org/web/20260702044745/https://signal-ai.com/about/) Pricing NOSIBLE See nosible.com [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Signal AI Contact Signal AI for access and commercial terms [[ 1 a]](https://archive.is/98DLH) [[ 1 b]](https://web.archive.org/web/20260517020355/https://signal-ai.com/solutions/api/) | Dimension | NOSIBLE | Signal AI | | --- | --- | --- | | Category | Web-scale search and market-event intelligence for AI agents [[ 4 a]](https://archive.is/glqrw) | Reputation and risk intelligence platform [[ 2 a]](https://archive.is/YQkKk) [[ 2 b]](https://web.archive.org/web/20260702044745/https://signal-ai.com/about/) | | Illustrative users | Investors, quants, AI agents [[ 4 a]](https://archive.is/glqrw) | Reputation, communications, ESG, supply-chain, and enterprise-risk teams [[ 2 a]](https://archive.is/YQkKk) [[ 2 b]](https://web.archive.org/web/20260702044745/https://signal-ai.com/about/) | | Content scope | Long-form news, corporate, government; excludes social, paywalled, marketplaces, adult [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Global news data including premium licensed content [[ 1 a]](https://archive.is/98DLH) [[ 1 b]](https://web.archive.org/web/20260517020355/https://signal-ai.com/solutions/api/) | | Languages | 95 [[ 4 a]](https://archive.is/glqrw) | Cited API page: 75+ covered and translated languages [[ 1 a]](https://archive.is/98DLH) [[ 1 b]](https://web.archive.org/web/20260517020355/https://signal-ai.com/solutions/api/) | | Published corpus indicators | 300,000+ open-web sources with source history [[ 4 a]](https://archive.is/glqrw) | Search across more than one billion documents [[ 1 a]](https://archive.is/98DLH) [[ 1 b]](https://web.archive.org/web/20260517020355/https://signal-ai.com/solutions/api/) | | Event database | Standalone 100M+ dated, ranked events (WORLD) [[ 5 a]](https://archive.is/tGYc9) | Events endpoint for monitoring and historical deep dives [[ 1 a]](https://archive.is/98DLH) [[ 1 b]](https://web.archive.org/web/20260517020355/https://signal-ai.com/solutions/api/) | | Analytical measures | Available in enrichment [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Metrics endpoint for trends, patterns, and themes [[ 1 a]](https://archive.is/98DLH) [[ 1 b]](https://web.archive.org/web/20260517020355/https://signal-ai.com/solutions/api/) | | Relationship and aspect analysis | Aspect-based sentiment is on the roadmap [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Affinity endpoint for relationships between concepts [[ 1 a]](https://archive.is/98DLH) [[ 1 b]](https://web.archive.org/web/20260517020355/https://signal-ai.com/solutions/api/) | | Entity and ticker mapping | Source-attributed entities [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Topic and concept metadata for reputation analysis [[ 2 a]](https://archive.is/YQkKk) [[ 2 b]](https://web.archive.org/web/20260702044745/https://signal-ai.com/about/) | | Delivery and API | Six-endpoint agent API, SDKs, MCP [[ 4 a]](https://archive.is/glqrw) | API access for BI solutions and AI agents [[ 2 a]](https://archive.is/YQkKk) [[ 2 b]](https://web.archive.org/web/20260702044745/https://signal-ai.com/about/) | | Agent-readiness | Built for AI agents [[ 4 a]](https://archive.is/glqrw) | Direct support for AI-agent integration [[ 1 a]](https://archive.is/98DLH) [[ 1 b]](https://web.archive.org/web/20260517020355/https://signal-ai.com/solutions/api/) | | Published use cases | Product materials describe event research and backtesting workflows [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Published use cases include reputation, ESG perception, and supply-chain risk [[ 2 a]](https://archive.is/YQkKk) [[ 2 b]](https://web.archive.org/web/20260702044745/https://signal-ai.com/about/) | | Pricing | See nosible.com [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Contact Signal AI for access and commercial terms [[ 1 a]](https://archive.is/98DLH) [[ 1 b]](https://web.archive.org/web/20260517020355/https://signal-ai.com/solutions/api/) | ## Potential fit by workflow Signal AI's cited materials describe reputation, communications, ESG-perception, and supply-chain workflows delivered through its platform and APIs. [[ 2 a]](https://archive.is/YQkKk) [[ 2 b]](https://web.archive.org/web/20260702044745/https://signal-ai.com/about/) NOSIBLE's cited materials describe dated open-web documents, ranked events, and ticker context for research agents. [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) In NOSIBLE's view, consider the product whose published workflow matches the requirement; the products may also be complementary. ## Common Signal AI comparison questions ### How does NOSIBLE feel it differentiates itself from Signal AI? NOSIBLE is an AI-native company with two products: SEARCH and WORLD. [[ 5 a]](https://archive.is/tGYc9) [[ 7 a]](https://archive.is/iDZeo) SEARCH lets agents find dated open-web sources they can cite and inspect directly. [[ 4 a]](https://archive.is/glqrw) WORLD is a live open-web event database for models and backtests, with an embedding per event. [[ 5 a]](https://archive.is/tGYc9) [[ 6 a]](https://archive.is/tnkpG) [[ 8 a]](https://archive.is/1CUvQ) NOSIBLE is committed to open-source software and makes its models publicly available on Hugging Face. [[ 9 a]](https://archive.is/7SSz0) [[ 10 a]](https://archive.is/kHxMG) Related [WORLD v1.2 trial](https://nosible.com/start-trial#data-coverage) [Sentiment model](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) [Forward-looking model](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) ### What if we need share of voice, readership, AI citations, and crisis dashboards? Signal AI publishes communications and reputation workflows. [[ 2 a]](https://archive.is/YQkKk) [[ 2 b]](https://web.archive.org/web/20260702044745/https://signal-ai.com/about/) NOSIBLE is aimed at dated event evidence, ticker mapping, and research-agent retrieval. [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) In NOSIBLE's view, Signal AI is the more direct fit for the first requirement. Buyers should choose the workflow that matches the primary user rather than expect one interface to replace every capability. Related [WORLD event database](https://nosible.world/world) ### How does Signal AI's licensed media, social, broadcast, podcast, and regulatory coverage compare with NOSIBLE? Signal AI publishes very broad global news coverage with premium licensing and multilingual translation. [[ 1 a]](https://archive.is/98DLH) [[ 1 b]](https://web.archive.org/web/20260517020355/https://signal-ai.com/solutions/api/) NOSIBLE uses a different, long-form open-web source model. [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Compare source access, rights, duplication, historical depth, and the ability to inspect original evidence on a representative query set. Related [Copyright posture](https://nosible.com/legal/copyright) [Security posture](https://nosible.com/legal/security) [WORLD event database](https://nosible.world/world) ### Can Signal AI's published AI-agent workflow replace a NOSIBLE-powered research agent? Signal AI's cited materials describe AI-agent integration for its reputation and risk data. [[ 2 a]](https://archive.is/YQkKk) [[ 2 b]](https://web.archive.org/web/20260702044745/https://signal-ai.com/about/) A NOSIBLE-powered agent is oriented toward dated source retrieval, event history, citations, and ticker context. [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) The products publish different workflow goals, so replacement should be evaluated task by task. [[ 1 a]](https://archive.is/98DLH) [[ 1 b]](https://web.archive.org/web/20260517020355/https://signal-ai.com/solutions/api/) [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Related [Agentic Search](https://docs.nosible.com/) [WORLD event database](https://nosible.world/world) ### What do Signal AI's API endpoints provide that NOSIBLE does not? Signal AI publishes Search, Metrics, Affinity, and Events endpoints for reputation intelligence. [[ 2 a]](https://archive.is/YQkKk) [[ 2 b]](https://web.archive.org/web/20260702044745/https://signal-ai.com/about/) NOSIBLE exposes search, source history, crawl, and ranked-event workflows with ticker and entity context. [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) The distinction is finished reputation analytics versus a cross-domain research evidence layer. [[ 2 a]](https://archive.is/YQkKk) [[ 2 b]](https://web.archive.org/web/20260702044745/https://signal-ai.com/about/) [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Related [Agentic Search](https://docs.nosible.com/) [WORLD event database](https://nosible.world/world) ### How does NOSIBLE support ESG or regulatory horizon scanning? Signal AI publishes ESG, regulatory, reputation, and supply-chain workflows. [[ 2 a]](https://archive.is/YQkKk) [[ 2 b]](https://web.archive.org/web/20260702044745/https://signal-ai.com/about/) NOSIBLE supports custom horizon scanning from dated evidence and ranked events. [[ 5 a]](https://archive.is/tGYc9) In NOSIBLE's view, consider Signal AI for the ready-made workflow or NOSIBLE when the team wants to construct and validate its own event model. Related [WORLD event database](https://nosible.world/world) [Semantic factors](https://nosible.com/semantic-factors) Continue comparing ## Related Vendor Comparisons Compare Signal AI with adjacent options for task execution, agent tooling, data access, and research workflows, then evaluate the output shape your application actually needs for production. [Opoint](https://nosible.com/compare/nosible-vs-opoint) [RepRisk](https://nosible.com/compare/nosible-vs-reprisk) [SESAMm](https://nosible.com/compare/nosible-vs-sesamm) [AlphaSense](https://nosible.com/compare/nosible-vs-alphasense) Dive deeper ## Take your next step today Review the delivered field definitions, classification boundaries, and example values before comparing vendor workflows, so your team can assess what each product actually returns to downstream systems. [Data dictionaries](https://nosible.com/data-dictionaries) [Ontology reference](https://nosible.com/ontologies) [API reference](https://docs.nosible.com/) [Start trial](https://nosible.com/start-trial) Sources reviewed July 14, 2026 : [[ 1 ] Signal AI API](https://signal-ai.com/solutions/api/) ( [dated snapshot](https://archive.is/98DLH); [Wayback copy](https://web.archive.org/web/20260517020355/https://signal-ai.com/solutions/api/) ) , [[ 2 ] Signal AI company and platform overview](https://signal-ai.com/about/) ( [dated snapshot](https://archive.is/YQkKk); [Wayback copy](https://web.archive.org/web/20260702044745/https://signal-ai.com/about/) ) , [[ 3 ] NOSIBLE product overview](https://nosible.com/) ( [dated snapshot](https://archive.is/r7qDv); [Wayback copy](https://web.archive.org/web/20260713185114/https://nosible.com/) ) , [[ 4 ] NOSIBLE SEARCH](https://docs.nosible.com/) ( [dated snapshot](https://archive.is/glqrw) ) , [[ 5 ] NOSIBLE WORLD](https://nosible.world/world) ( [dated snapshot](https://archive.is/tGYc9) ) , [[ 6 ] NOSIBLE WORLD v1.2 trial and coverage](https://nosible.com/start-trial) ( [dated snapshot](https://archive.is/tnkpG) ) , [[ 7 ] NOSIBLE AI-native research overview](https://nosible.com/blog) ( [dated snapshot](https://archive.is/iDZeo) ) , [[ 8 ] NOSIBLE embedding-based research](https://nosible.com/blog/an-embedding-based-approach-to-trade-and-economic-policy-uncertainty) ( [dated snapshot](https://archive.is/1CUvQ) ) , [[ 9 ] NOSIBLE Financial Sentiment v1.2 Base](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) ( [dated snapshot](https://archive.is/7SSz0) ) , [[ 10 ] NOSIBLE Forward-Looking v1.2 Base](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) ( [dated snapshot](https://archive.is/kHxMG) ) . This page compares NOSIBLE WORLD with Signal AI. Signal AI and related product names are marks of their respective owner. NOSIBLE is not affiliated with, sponsored by, or endorsed by Signal AI. Comparison standards, legal context & corrections This comparison was prepared by NOSIBLE, which has a commercial interest in the products being compared. It is based on the cited public materials as they appeared on July 14, 2026. NOSIBLE has not tested every competitor feature, and the page is not a complete statement of either product. Products change; confirm current requirements, availability, and commercial terms with each vendor. Factual claims are attributed to cited first-party materials; evaluative statements reflect NOSIBLE's opinion. Each source includes a dated archive.is snapshot and, where the Internet Archive captured that URL, a timestamp-specific Wayback copy. Archive availability is controlled by those services. Competitor names and marks are used only to identify the products being compared. No affiliation, sponsorship, or endorsement is implied. The FTC says truthful, non-deceptive comparative advertising may identify competitors ( [policy](https://www.ftc.gov/legal-library/browse/statement-policy-regarding-comparative-advertising); [dated snapshot](https://archive.is/nZXY4); [Wayback copy](https://web.archive.org/web/20260215042003/https://www.ftc.gov/legal-library/browse/statement-policy-regarding-comparative-advertising) ). For one U.S. example of nominative-use analysis, see *New Kids on the Block v. News America Publishing, Inc.*, 971 F.2d 302, 308 (9th Cir. 1992) ( [opinion](https://law.justia.com/cases/federal/appellate-courts/F2/971/302/72076/); [dated snapshot](https://archive.is/dnk8j); [Wayback copy](https://web.archive.org/web/20250211213616/https://law.justia.com/cases/federal/appellate-courts/F2/971/302/72076/) ). If you represent Signal AI and believe a factual statement is inaccurate, email [stuart@nosible.com](mailto:stuart@nosible.com?subject=Correction%20request%3A%20NOSIBLE%20WORLD%20vs%20Signal%20AI) with the specific claim and a supporting first-party URL. NOSIBLE will review and correct substantiated errors. > How NOSIBLE WORLD, a web-scale search and market-event intelligence engine for AI agents, compares with Signal AI, a reputation and risk intelligence platform. Coverage, point-in-time integrity, scope, and use cases. **URL:** https://nosible.com/compare/nosible-vs-signal-ai --- --- title: "NOSIBLE vs MarketPsych" description: "Compare NOSIBLE and MarketPsych for sentiment analytics, AI agents, point-in-time search, event intelligence, language coverage, and source retrieval." url: "https://nosible.com/compare/nosible-vs-marketpsych" --- Comparison / Reviewed July 14, 2026 # NOSIBLE vs MarketPsych LSEG MarketPsych turns curated news and social content into real-time sentiment scores. [[ 2 a]](https://archive.is/Qu2EM) [[ 2 b]](https://web.archive.org/web/20240918140802/https://www.lseg.com/content/dam/marketing/en_us/documents/white-papers/refinitiv-marketpsych-esg-analytics-whitepaper.pdf) NOSIBLE gives agents dated documents and ranked events through an open-web workflow. [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) NOSIBLE-AUTHORED COMPARISON · NOSIBLE HAS A COMMERCIAL INTEREST IN THIS COMPARISON · FACTS ATTRIBUTED TO FIRST-PARTY VENDOR MATERIALS · EVALUATIVE STATEMENTS ARE NOSIBLE'S OPINION · REVIEWED JULY 14, 2026 · DUAL PUBLIC ARCHIVES WHERE SUPPORTED · [STANDARDS & CORRECTIONS](https://nosible.com/compare/nosible-vs-marketpsych#comparison-standards). IF YOU REPRESENT MARKETPSYCH AND BELIEVE A FACTUAL STATEMENT IS INACCURATE, EMAIL [STUART@NOSIBLE.COM](mailto:stuart@nosible.com?subject=Correction%20request%3A%20NOSIBLE%20vs%20MarketPsych) WITH THE SPECIFIC CLAIM AND A SUPPORTING FIRST-PARTY URL. NOSIBLE WILL REVIEW AND CORRECT SUBSTANTIATED ERRORS. - LSEG's cited September 2023 MarketPsych fact sheet says it converts 4,000+ news and social outlets into more than 100 sentiment scores. [[ 2 a]](https://archive.is/Qu2EM) [[ 2 b]](https://web.archive.org/web/20240918140802/https://www.lseg.com/content/dam/marketing/en_us/documents/white-papers/refinitiv-marketpsych-esg-analytics-whitepaper.pdf) - The published dataset covers 12 languages, history from 1998, and minute, hourly, and daily updates. [[ 2 a]](https://archive.is/Qu2EM) [[ 2 b]](https://web.archive.org/web/20240918140802/https://www.lseg.com/content/dam/marketing/en_us/documents/white-papers/refinitiv-marketpsych-esg-analytics-whitepaper.pdf) - Its reported universe includes companies, countries, indexes, currencies, commodities, and cryptocurrencies. [[ 1 a]](https://archive.is/fds1x) [[ 1 b]](https://web.archive.org/web/20260713190929/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/lseg-marketpsych-analytics-factsheet.pdf) - NOSIBLE focuses on dated source retrieval, ranked events, and agent workflows. [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) - In NOSIBLE's view, consider NOSIBLE when you need inspectable documents and events rather than a finished sentiment series. ## Sentiment scores and source-level evidence MarketPsych is a cross-asset sentiment dataset with published entity coverage, a long time series, and frequently updated scores. [[ 1 a]](https://archive.is/fds1x) [[ 1 b]](https://web.archive.org/web/20260713190929/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/lseg-marketpsych-analytics-factsheet.pdf) NOSIBLE lets agents and researchers search dated source material, retrieve ranked events, and build custom workflows from open-web evidence. [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) In NOSIBLE's view, NOSIBLE is the stronger fit when the system must inspect and cite the documents behind a conclusion. ## Sentiment analytics and source-level evidence The competitor column summarizes what the named vendor's cited first-party materials describe; the NOSIBLE column summarizes NOSIBLE's cited materials. Primary use NOSIBLE Search and event intelligence [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) MarketPsych Sentiment and behavioral analytics [[ 1 a]](https://archive.is/fds1x) [[ 1 b]](https://web.archive.org/web/20260713190929/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/lseg-marketpsych-analytics-factsheet.pdf) Illustrative users NOSIBLE AI agents, backtests, risk systems [[ 4 a]](https://archive.is/glqrw) MarketPsych Quant, research, risk, and macro teams [[ 1 a]](https://archive.is/fds1x) [[ 1 b]](https://web.archive.org/web/20260713190929/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/lseg-marketpsych-analytics-factsheet.pdf) Primary output NOSIBLE Dated documents and ranked events [[ 5 a]](https://archive.is/tGYc9) MarketPsych 100+ structured sentiment and topic scores [[ 1 a]](https://archive.is/fds1x) [[ 1 b]](https://web.archive.org/web/20260713190929/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/lseg-marketpsych-analytics-factsheet.pdf) Content scope NOSIBLE News, corporate, government text [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) MarketPsych Cited September 2023 fact sheet: 4,000+ news and social-media outlets [[ 2 a]](https://archive.is/Qu2EM) [[ 2 b]](https://web.archive.org/web/20240918140802/https://www.lseg.com/content/dam/marketing/en_us/documents/white-papers/refinitiv-marketpsych-esg-analytics-whitepaper.pdf) Languages NOSIBLE 95 [[ 4 a]](https://archive.is/glqrw) MarketPsych 12 [[ 1 a]](https://archive.is/fds1x) [[ 1 b]](https://web.archive.org/web/20260713190929/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/lseg-marketpsych-analytics-factsheet.pdf) History NOSIBLE Roughly 30 years [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) MarketPsych 1998 to present [[ 2 a]](https://archive.is/Qu2EM) [[ 2 b]](https://web.archive.org/web/20240918140802/https://www.lseg.com/content/dam/marketing/en_us/documents/white-papers/refinitiv-marketpsych-esg-analytics-whitepaper.pdf) Point-in-time method NOSIBLE Five-way date verification [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) MarketPsych Historical time series with real-time updates [[ 1 a]](https://archive.is/fds1x) [[ 1 b]](https://web.archive.org/web/20260713190929/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/lseg-marketpsych-analytics-factsheet.pdf) Primary data model NOSIBLE Ranked dated events [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) MarketPsych Sentiment, emotion, and fundamental-theme measures [[ 1 a]](https://archive.is/fds1x) [[ 1 b]](https://web.archive.org/web/20260713190929/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/lseg-marketpsych-analytics-factsheet.pdf) Delivery NOSIBLE API, SDKs, MCP, agentic search [[ 4 a]](https://archive.is/glqrw) MarketPsych 60-second, hourly, and daily data [[ 1 a]](https://archive.is/fds1x) [[ 1 b]](https://web.archive.org/web/20260713190929/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/lseg-marketpsych-analytics-factsheet.pdf) Agent readiness NOSIBLE Built for AI agents [[ 4 a]](https://archive.is/glqrw) MarketPsych Structured data for dashboards, statistical tools, and models [[ 1 a]](https://archive.is/fds1x) [[ 1 b]](https://web.archive.org/web/20260713190929/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/lseg-marketpsych-analytics-factsheet.pdf) | Dimension | NOSIBLE | MarketPsych | | --- | --- | --- | | Primary use | Search and event intelligence [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Sentiment and behavioral analytics [[ 1 a]](https://archive.is/fds1x) [[ 1 b]](https://web.archive.org/web/20260713190929/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/lseg-marketpsych-analytics-factsheet.pdf) | | Illustrative users | AI agents, backtests, risk systems [[ 4 a]](https://archive.is/glqrw) | Quant, research, risk, and macro teams [[ 1 a]](https://archive.is/fds1x) [[ 1 b]](https://web.archive.org/web/20260713190929/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/lseg-marketpsych-analytics-factsheet.pdf) | | Primary output | Dated documents and ranked events [[ 5 a]](https://archive.is/tGYc9) | 100+ structured sentiment and topic scores [[ 1 a]](https://archive.is/fds1x) [[ 1 b]](https://web.archive.org/web/20260713190929/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/lseg-marketpsych-analytics-factsheet.pdf) | | Content scope | News, corporate, government text [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Cited September 2023 fact sheet: 4,000+ news and social-media outlets [[ 2 a]](https://archive.is/Qu2EM) [[ 2 b]](https://web.archive.org/web/20240918140802/https://www.lseg.com/content/dam/marketing/en_us/documents/white-papers/refinitiv-marketpsych-esg-analytics-whitepaper.pdf) | | Languages | 95 [[ 4 a]](https://archive.is/glqrw) | 12 [[ 1 a]](https://archive.is/fds1x) [[ 1 b]](https://web.archive.org/web/20260713190929/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/lseg-marketpsych-analytics-factsheet.pdf) | | History | Roughly 30 years [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | 1998 to present [[ 2 a]](https://archive.is/Qu2EM) [[ 2 b]](https://web.archive.org/web/20240918140802/https://www.lseg.com/content/dam/marketing/en_us/documents/white-papers/refinitiv-marketpsych-esg-analytics-whitepaper.pdf) | | Point-in-time method | Five-way date verification [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Historical time series with real-time updates [[ 1 a]](https://archive.is/fds1x) [[ 1 b]](https://web.archive.org/web/20260713190929/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/lseg-marketpsych-analytics-factsheet.pdf) | | Primary data model | Ranked dated events [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Sentiment, emotion, and fundamental-theme measures [[ 1 a]](https://archive.is/fds1x) [[ 1 b]](https://web.archive.org/web/20260713190929/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/lseg-marketpsych-analytics-factsheet.pdf) | | Delivery | API, SDKs, MCP, agentic search [[ 4 a]](https://archive.is/glqrw) | 60-second, hourly, and daily data [[ 1 a]](https://archive.is/fds1x) [[ 1 b]](https://web.archive.org/web/20260713190929/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/lseg-marketpsych-analytics-factsheet.pdf) | | Agent readiness | Built for AI agents [[ 4 a]](https://archive.is/glqrw) | Structured data for dashboards, statistical tools, and models [[ 1 a]](https://archive.is/fds1x) [[ 1 b]](https://web.archive.org/web/20260713190929/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/lseg-marketpsych-analytics-factsheet.pdf) | ## Potential fit by workflow MarketPsych's cited materials describe ready-made cross-asset sentiment time series. [[ 1 a]](https://archive.is/fds1x) [[ 1 b]](https://web.archive.org/web/20260713190929/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/lseg-marketpsych-analytics-factsheet.pdf) NOSIBLE's cited materials describe dated documents, ranked events, and inspectable source context for agents. [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) In NOSIBLE's view, consider the product whose published output matches the requirement; teams may also use NOSIBLE evidence to investigate a MarketPsych signal. ## Common MarketPsych comparison questions ### How does NOSIBLE feel it differentiates itself from MarketPsych? NOSIBLE is an AI-native company with two products: SEARCH and WORLD. [[ 5 a]](https://archive.is/tGYc9) [[ 7 a]](https://archive.is/iDZeo) SEARCH lets agents find dated open-web sources they can cite and inspect directly. [[ 4 a]](https://archive.is/glqrw) WORLD is a live open-web event database for models and backtests, with an embedding per event. [[ 5 a]](https://archive.is/tGYc9) [[ 6 a]](https://archive.is/tnkpG) [[ 8 a]](https://archive.is/1CUvQ) NOSIBLE is committed to open-source software and makes its models publicly available on Hugging Face. [[ 9 a]](https://archive.is/7SSz0) [[ 10 a]](https://archive.is/kHxMG) Related [WORLD v1.2 trial](https://nosible.com/start-trial#data-coverage) [Sentiment model](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) [Forward-looking model](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) ### Do we need MarketPsych's ready-made sentiment indices or NOSIBLE's source-level evidence? MarketPsych publishes ready-made sentiment series; NOSIBLE publishes source-retrieval and event workflows. [[ 1 a]](https://archive.is/fds1x) [[ 1 b]](https://web.archive.org/web/20260713190929/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/lseg-marketpsych-analytics-factsheet.pdf) [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) In NOSIBLE's view, MarketPsych is the more direct choice for the first requirement, while NOSIBLE is the better fit for the second. Compare both against the actual downstream model rather than treating one as a numerical substitute for the other. Related [Sentiment model](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) [Sentiment research](https://nosible.com/blog/fast-enough-to-matter-productionizing-tiny-transformers-for-signal-extraction) [WORLD event database](https://nosible.world/world) ### How important are social-media buzz and author or channel signals to our strategy? MarketPsych expressly includes thousands of social channels and measures buzz, emotion, and sentiment. [[ 1 a]](https://archive.is/fds1x) [[ 1 b]](https://web.archive.org/web/20260713190929/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/lseg-marketpsych-analytics-factsheet.pdf) NOSIBLE emphasizes long-form open-web evidence instead. [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) In NOSIBLE's view, strategies that depend on social attention may prefer MarketPsych, while strategies that require inspectable documents and event context may prefer NOSIBLE. Related [Security posture](https://nosible.com/legal/security) [Sentiment model](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) [Sentiment research](https://nosible.com/blog/fast-enough-to-matter-productionizing-tiny-transformers-for-signal-extraction) ### Can NOSIBLE replace MarketPsych for currencies, commodities, sovereigns, and crypto sentiment? Not as a drop-in product. [[ 1 a]](https://archive.is/fds1x) [[ 1 b]](https://web.archive.org/web/20260713190929/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/lseg-marketpsych-analytics-factsheet.pdf) [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) MarketPsych publishes cross-asset coverage and finished scores. [[ 1 a]](https://archive.is/fds1x) [[ 1 b]](https://web.archive.org/web/20260713190929/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/lseg-marketpsych-analytics-factsheet.pdf) NOSIBLE supplies documents, events, entity mappings, and enrichment tools from which a team can design a proprietary indicator. [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) That approach requires additional modelling and validation. [[ 1 a]](https://archive.is/fds1x) [[ 1 b]](https://web.archive.org/web/20260713190929/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/lseg-marketpsych-analytics-factsheet.pdf) [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Related [Sentiment model](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) [Sentiment research](https://nosible.com/blog/fast-enough-to-matter-productionizing-tiny-transformers-for-signal-extraction) [WORLD event database](https://nosible.world/world) ### How should we validate MarketPsych's long history against NOSIBLE's point-in-time claims? Define a common asset universe, timestamp convention, source period, and target outcome. MarketPsych publishes history from 1998; NOSIBLE emphasizes replayable dated sources and events. [[ 2 a]](https://archive.is/Qu2EM) [[ 2 b]](https://web.archive.org/web/20240918140802/https://www.lseg.com/content/dam/marketing/en_us/documents/white-papers/refinitiv-marketpsych-esg-analytics-whitepaper.pdf) [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Evaluate stability, revisions, missing data, and predictive performance separately for each product. Related [WORLD event database](https://nosible.world/world) ### What should an LLM agent do with MarketPsych scores versus NOSIBLE documents? A MarketPsych score can serve as a structured feature. [[ 1 a]](https://archive.is/fds1x) [[ 1 b]](https://web.archive.org/web/20260713190929/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/lseg-marketpsych-analytics-factsheet.pdf) NOSIBLE can supply the documents and events an agent uses for explanation, retrieval, and follow-up analysis. [[ 4 a]](https://archive.is/glqrw) A combined system should preserve each vendor's timestamp, identifier, and source lineage. Related [Agentic Search](https://docs.nosible.com/) [WORLD event database](https://nosible.world/world) Continue comparing ## Related Vendor Comparisons Compare MarketPsych with adjacent options for task execution, agent tooling, data access, and research workflows, then evaluate the output shape your application actually needs for production. [RavenPack](https://nosible.com/compare/nosible-vs-ravenpack) [LSEG](https://nosible.com/compare/nosible-vs-lseg) [MKT MediaStats](https://nosible.com/compare/nosible-vs-mkt-mediastats) [GDELT](https://nosible.com/compare/nosible-vs-gdelt) Dive deeper ## Take your next step today Review the delivered field definitions, classification boundaries, and example values before comparing vendor workflows, so your team can assess what each product actually returns to downstream systems. [Data dictionaries](https://nosible.com/data-dictionaries) [Ontology reference](https://nosible.com/ontologies) [API reference](https://docs.nosible.com/) [Start trial](https://nosible.com/start-trial) Sources reviewed July 14, 2026 : [[ 1 ] LSEG MarketPsych factsheet](https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/lseg-marketpsych-analytics-factsheet.pdf) ( [dated snapshot](https://archive.is/fds1x); [Wayback copy](https://web.archive.org/web/20260713190929/https://www.lseg.com/content/dam/data-analytics/en_us/documents/fact-sheets/lseg-marketpsych-analytics-factsheet.pdf) ) , [[ 2 ] LSEG MarketPsych ESG methodology](https://www.lseg.com/content/dam/marketing/en_us/documents/white-papers/refinitiv-marketpsych-esg-analytics-whitepaper.pdf) ( [dated snapshot](https://archive.is/Qu2EM); [Wayback copy](https://web.archive.org/web/20240918140802/https://www.lseg.com/content/dam/marketing/en_us/documents/white-papers/refinitiv-marketpsych-esg-analytics-whitepaper.pdf) ) , [[ 3 ] NOSIBLE product overview](https://nosible.com/) ( [dated snapshot](https://archive.is/r7qDv); [Wayback copy](https://web.archive.org/web/20260713185114/https://nosible.com/) ) , [[ 4 ] NOSIBLE SEARCH](https://docs.nosible.com/) ( [dated snapshot](https://archive.is/glqrw) ) , [[ 5 ] NOSIBLE WORLD](https://nosible.world/world) ( [dated snapshot](https://archive.is/tGYc9) ) , [[ 6 ] NOSIBLE WORLD v1.2 trial and coverage](https://nosible.com/start-trial) ( [dated snapshot](https://archive.is/tnkpG) ) , [[ 7 ] NOSIBLE AI-native research overview](https://nosible.com/blog) ( [dated snapshot](https://archive.is/iDZeo) ) , [[ 8 ] NOSIBLE embedding-based research](https://nosible.com/blog/an-embedding-based-approach-to-trade-and-economic-policy-uncertainty) ( [dated snapshot](https://archive.is/1CUvQ) ) , [[ 9 ] NOSIBLE Financial Sentiment v1.2 Base](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) ( [dated snapshot](https://archive.is/7SSz0) ) , [[ 10 ] NOSIBLE Forward-Looking v1.2 Base](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) ( [dated snapshot](https://archive.is/kHxMG) ) . MarketPsych and LSEG are trademarks of their respective owners. NOSIBLE is not affiliated with, sponsored by, or endorsed by MarketPsych or LSEG. Comparison standards, legal context & corrections This comparison was prepared by NOSIBLE, which has a commercial interest in the products being compared. It is based on the cited public materials as they appeared on July 14, 2026. NOSIBLE has not tested every competitor feature, and the page is not a complete statement of either product. Products change; confirm current requirements, availability, and commercial terms with each vendor. Factual claims are attributed to cited first-party materials; evaluative statements reflect NOSIBLE's opinion. Each source includes a dated archive.is snapshot and, where the Internet Archive captured that URL, a timestamp-specific Wayback copy. Archive availability is controlled by those services. Competitor names and marks are used only to identify the products being compared. No affiliation, sponsorship, or endorsement is implied. The FTC says truthful, non-deceptive comparative advertising may identify competitors ( [policy](https://www.ftc.gov/legal-library/browse/statement-policy-regarding-comparative-advertising); [dated snapshot](https://archive.is/nZXY4); [Wayback copy](https://web.archive.org/web/20260215042003/https://www.ftc.gov/legal-library/browse/statement-policy-regarding-comparative-advertising) ). For one U.S. example of nominative-use analysis, see *New Kids on the Block v. News America Publishing, Inc.*, 971 F.2d 302, 308 (9th Cir. 1992) ( [opinion](https://law.justia.com/cases/federal/appellate-courts/F2/971/302/72076/); [dated snapshot](https://archive.is/dnk8j); [Wayback copy](https://web.archive.org/web/20250211213616/https://law.justia.com/cases/federal/appellate-courts/F2/971/302/72076/) ). If you represent MarketPsych and believe a factual statement is inaccurate, email [stuart@nosible.com](mailto:stuart@nosible.com?subject=Correction%20request%3A%20NOSIBLE%20vs%20MarketPsych) with the specific claim and a supporting first-party URL. NOSIBLE will review and correct substantiated errors. > Compare NOSIBLE and MarketPsych for sentiment analytics, AI agents, point-in-time search, event intelligence, language coverage, and source retrieval. **URL:** https://nosible.com/compare/nosible-vs-marketpsych --- --- title: "NOSIBLE vs MKT MediaStats" description: "Compare NOSIBLE and MKT MediaStats for point-in-time event intelligence, raw source retrieval, AI agents, factor workflows, delivery, and backtesting." url: "https://nosible.com/compare/nosible-vs-mkt-mediastats" --- Comparison / Reviewed July 14, 2026 # NOSIBLE vs MKT MediaStats MKT MediaStats converts media and behavioral data into narrative indicators, indexes, and institutional research. [[ 1 a]](https://archive.is/85YK0) [[ 1 b]](https://web.archive.org/web/20260510170725/https://www.mktmediastats.com/) NOSIBLE supplies dated source material and ranked events for custom agent and quant workflows. [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) NOSIBLE-AUTHORED COMPARISON · NOSIBLE HAS A COMMERCIAL INTEREST IN THIS COMPARISON · FACTS ATTRIBUTED TO FIRST-PARTY VENDOR MATERIALS · EVALUATIVE STATEMENTS ARE NOSIBLE'S OPINION · REVIEWED JULY 14, 2026 · DUAL PUBLIC ARCHIVES WHERE SUPPORTED · [STANDARDS & CORRECTIONS](https://nosible.com/compare/nosible-vs-mkt-mediastats#comparison-standards). IF YOU REPRESENT MKT MEDIASTATS AND BELIEVE A FACTUAL STATEMENT IS INACCURATE, EMAIL [STUART@NOSIBLE.COM](mailto:stuart@nosible.com?subject=Correction%20request%3A%20NOSIBLE%20vs%20MKT%20MediaStats) WITH THE SPECIFIC CLAIM AND A SUPPORTING FIRST-PARTY URL. NOSIBLE WILL REVIEW AND CORRECT SUBSTANTIATED ERRORS. - MKT MediaStats says it analyzes millions of media and behavioral channels to identify financial and economic narratives. [[ 1 a]](https://archive.is/85YK0) [[ 1 b]](https://web.archive.org/web/20260510170725/https://www.mktmediastats.com/) - Its offerings include alpha indicators, market strategy, thematic baskets, political analysis, risk assessment, and advisory services. [[ 1 a]](https://archive.is/85YK0) [[ 1 b]](https://web.archive.org/web/20260510170725/https://www.mktmediastats.com/) - MKT MediaStats identifies partnerships with State Street Global Markets and MSCI for indicators and indexes. [[ 1 a]](https://archive.is/85YK0) [[ 1 b]](https://web.archive.org/web/20260510170725/https://www.mktmediastats.com/) - An official MKT MediaStats research note describes a multi-asset strategy based on narrative momentum. [[ 2 a]](https://archive.is/EI2sN) [[ 2 b]](https://web.archive.org/web/20260416182645/https://www.mktmediastats.com/post/multi-asset-rotation) - NOSIBLE emphasizes dated source retrieval, ranked events, and agent-oriented research. [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) - In NOSIBLE's view, consider NOSIBLE when you want to build and explain proprietary signals from source material. ## Narrative indicators and a source-and-event layer MKT MediaStats is designed to quantify financial and economic narratives and connect them to assets, portfolios, and operational outcomes. [[ 1 a]](https://archive.is/85YK0) [[ 1 b]](https://web.archive.org/web/20260510170725/https://www.mktmediastats.com/) NOSIBLE gives researchers and AI agents dated source retrieval, ranked events, and source attribution for building their own workflows. [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) In NOSIBLE's view, NOSIBLE is the stronger fit when proprietary evidence and model design are the priority. ## Narrative analytics and source-level event research The competitor column summarizes what the named vendor's cited first-party materials describe; the NOSIBLE column summarizes NOSIBLE's cited materials. Primary use NOSIBLE Point-in-time event intelligence [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) MKT MediaStats Financial and economic narrative analytics [[ 1 a]](https://archive.is/85YK0) [[ 1 b]](https://web.archive.org/web/20260510170725/https://www.mktmediastats.com/) Illustrative users NOSIBLE AI agents, quant research, risk systems [[ 4 a]](https://archive.is/glqrw) MKT MediaStats Portfolio, risk, strategy, and economic-research teams [[ 1 a]](https://archive.is/85YK0) [[ 1 b]](https://web.archive.org/web/20260510170725/https://www.mktmediastats.com/) Content scope NOSIBLE Long-form news, corporate, and government text [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) MKT MediaStats Media and behavioral data [[ 1 a]](https://archive.is/85YK0) [[ 1 b]](https://web.archive.org/web/20260510170725/https://www.mktmediastats.com/) Data shape NOSIBLE Raw searchable source material and events [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) MKT MediaStats Indicators, narrative sensitivities, baskets, and indexes [[ 1 a]](https://archive.is/85YK0) [[ 1 b]](https://web.archive.org/web/20260510170725/https://www.mktmediastats.com/) Source footprint NOSIBLE 300,000+ open-web sources [[ 4 a]](https://archive.is/glqrw) MKT MediaStats Millions of media and behavioral channels [[ 1 a]](https://archive.is/85YK0) [[ 1 b]](https://web.archive.org/web/20260510170725/https://www.mktmediastats.com/) Processing model NOSIBLE Five-way date verification for point-in-time retrieval [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) MKT MediaStats Systematic media and behavioral-data processing [[ 1 a]](https://archive.is/85YK0) [[ 1 b]](https://web.archive.org/web/20260510170725/https://www.mktmediastats.com/) Events NOSIBLE 100M+ ranked dated events [[ 5 a]](https://archive.is/tGYc9) MKT MediaStats Narratives and security-level sensitivities [[ 1 a]](https://archive.is/85YK0) [[ 1 b]](https://web.archive.org/web/20260510170725/https://www.mktmediastats.com/) Delivery NOSIBLE API, SDKs, MCP, Cybernaut-1 [[ 4 a]](https://archive.is/glqrw) MKT MediaStats Indicators, indexes, baskets, strategy, and advisory offerings [[ 1 a]](https://archive.is/85YK0) [[ 1 b]](https://web.archive.org/web/20260510170725/https://www.mktmediastats.com/) Pricing NOSIBLE Product access through NOSIBLE [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) MKT MediaStats Contact MKT MediaStats for access and fees [[ 1 a]](https://archive.is/85YK0) [[ 1 b]](https://web.archive.org/web/20260510170725/https://www.mktmediastats.com/) | Dimension | NOSIBLE | MKT MediaStats | | --- | --- | --- | | Primary use | Point-in-time event intelligence [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Financial and economic narrative analytics [[ 1 a]](https://archive.is/85YK0) [[ 1 b]](https://web.archive.org/web/20260510170725/https://www.mktmediastats.com/) | | Illustrative users | AI agents, quant research, risk systems [[ 4 a]](https://archive.is/glqrw) | Portfolio, risk, strategy, and economic-research teams [[ 1 a]](https://archive.is/85YK0) [[ 1 b]](https://web.archive.org/web/20260510170725/https://www.mktmediastats.com/) | | Content scope | Long-form news, corporate, and government text [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Media and behavioral data [[ 1 a]](https://archive.is/85YK0) [[ 1 b]](https://web.archive.org/web/20260510170725/https://www.mktmediastats.com/) | | Data shape | Raw searchable source material and events [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Indicators, narrative sensitivities, baskets, and indexes [[ 1 a]](https://archive.is/85YK0) [[ 1 b]](https://web.archive.org/web/20260510170725/https://www.mktmediastats.com/) | | Source footprint | 300,000+ open-web sources [[ 4 a]](https://archive.is/glqrw) | Millions of media and behavioral channels [[ 1 a]](https://archive.is/85YK0) [[ 1 b]](https://web.archive.org/web/20260510170725/https://www.mktmediastats.com/) | | Processing model | Five-way date verification for point-in-time retrieval [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Systematic media and behavioral-data processing [[ 1 a]](https://archive.is/85YK0) [[ 1 b]](https://web.archive.org/web/20260510170725/https://www.mktmediastats.com/) | | Events | 100M+ ranked dated events [[ 5 a]](https://archive.is/tGYc9) | Narratives and security-level sensitivities [[ 1 a]](https://archive.is/85YK0) [[ 1 b]](https://web.archive.org/web/20260510170725/https://www.mktmediastats.com/) | | Delivery | API, SDKs, MCP, Cybernaut-1 [[ 4 a]](https://archive.is/glqrw) | Indicators, indexes, baskets, strategy, and advisory offerings [[ 1 a]](https://archive.is/85YK0) [[ 1 b]](https://web.archive.org/web/20260510170725/https://www.mktmediastats.com/) | | Pricing | Product access through NOSIBLE [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Contact MKT MediaStats for access and fees [[ 1 a]](https://archive.is/85YK0) [[ 1 b]](https://web.archive.org/web/20260510170725/https://www.mktmediastats.com/) | ## Potential fit by workflow MKT MediaStats' cited materials describe narrative indicators, thematic baskets, and institutional strategy services. [[ 1 a]](https://archive.is/85YK0) [[ 1 b]](https://web.archive.org/web/20260510170725/https://www.mktmediastats.com/) NOSIBLE's cited materials describe dated source material and event history for agent and quantitative research. [[ 4 a]](https://archive.is/glqrw) In NOSIBLE's view, consider the product whose published output matches the requirement; a packaged indicator may also be used alongside NOSIBLE evidence. ## Common MKT MediaStats comparison questions ### How does NOSIBLE feel it differentiates itself from MKT MediaStats? NOSIBLE is an AI-native company with two products: SEARCH and WORLD. [[ 5 a]](https://archive.is/tGYc9) [[ 7 a]](https://archive.is/iDZeo) SEARCH lets agents find dated open-web sources they can cite and inspect directly. [[ 4 a]](https://archive.is/glqrw) WORLD is a live open-web event database for models and backtests, with an embedding per event. [[ 5 a]](https://archive.is/tGYc9) [[ 6 a]](https://archive.is/tnkpG) [[ 8 a]](https://archive.is/1CUvQ) NOSIBLE is committed to open-source software and makes its models publicly available on Hugging Face. [[ 9 a]](https://archive.is/7SSz0) [[ 10 a]](https://archive.is/kHxMG) Related [WORLD v1.2 trial](https://nosible.com/start-trial#data-coverage) [Sentiment model](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) [Forward-looking model](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) ### Do we need finished MediaStats indicators or raw source retrieval behind the signal? MKT MediaStats publishes finished narrative products; NOSIBLE publishes dated-source and event workflows for custom research. [[ 1 a]](https://archive.is/85YK0) [[ 1 b]](https://web.archive.org/web/20260510170725/https://www.mktmediastats.com/) [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) In NOSIBLE's view, MKT MediaStats is the more direct choice for the first requirement, while NOSIBLE is the better fit for the second. The custom approach requires more internal research and validation. [[ 1 a]](https://archive.is/85YK0) [[ 1 b]](https://web.archive.org/web/20260510170725/https://www.mktmediastats.com/) [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Related [WORLD event database](https://nosible.world/world) ### Can we audit which articles drove a MediaStats score? Ask MKT MediaStats what source-level lineage is available for the licensed product under consideration. NOSIBLE exposes its own dated documents and event context for inspection. [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Auditability depends on the exact dataset and contract, so it should be tested rather than inferred from the product category. [[ 1 a]](https://archive.is/85YK0) [[ 1 b]](https://web.archive.org/web/20260510170725/https://www.mktmediastats.com/) [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Related [WORLD event database](https://nosible.world/world) [Bulk Web Search](https://docs.nosible.com/endpoints/search-bulk) ### Is our investment universe covered by MediaStats? MKT MediaStats describes products across global stocks, sectors, country equity, foreign exchange, fixed income, cryptocurrencies, macro themes, and operational risks. [[ 1 a]](https://archive.is/85YK0) [[ 1 b]](https://web.archive.org/web/20260510170725/https://www.mktmediastats.com/) NOSIBLE starts from documents and events rather than a finished factor universe. [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Validate both against the securities and themes actually traded. Related [WORLD event database](https://nosible.world/world) [Geopolitical risk index](https://nosible.com/blog/rebuilding-the-geopolitical-risk-index-from-nosible-world) [Stock Screening](https://nosible.com/#proof) ### How are Media Linkages different from NOSIBLE event and entity retrieval? MKT MediaStats models narrative sensitivities and relationships as a finished analytical output. [[ 1 a]](https://archive.is/85YK0) [[ 1 b]](https://web.archive.org/web/20260510170725/https://www.mktmediastats.com/) NOSIBLE lets teams form their own relationship hypotheses from dated documents, entities, products, tickers, and events. [[ 5 a]](https://archive.is/tGYc9) They are different stages of a research process. [[ 1 a]](https://archive.is/85YK0) [[ 1 b]](https://web.archive.org/web/20260510170725/https://www.mktmediastats.com/) [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Related [WORLD event database](https://nosible.world/world) ### Are we buying an MSCI or MediaStats allocation index, or building proprietary signals? In NOSIBLE's view, consider the index or indicator when its defined methodology and delivery match the mandate; consider NOSIBLE when the objective is proprietary research from dated evidence and ranked events. Teams may also combine a licensed index with source-level research, subject to each vendor's terms. Related [WORLD event database](https://nosible.world/world) Continue comparing ## Related Vendor Comparisons Compare MKT MediaStats with adjacent options for task execution, agent tooling, data access, and research workflows, then evaluate the output shape your application actually needs for production. [MarketPsych](https://nosible.com/compare/nosible-vs-marketpsych) [RavenPack](https://nosible.com/compare/nosible-vs-ravenpack) [GDELT](https://nosible.com/compare/nosible-vs-gdelt) [LSEG](https://nosible.com/compare/nosible-vs-lseg) Dive deeper ## Take your next step today Review the delivered field definitions, classification boundaries, and example values before comparing vendor workflows, so your team can assess what each product actually returns to downstream systems. [Data dictionaries](https://nosible.com/data-dictionaries) [Ontology reference](https://nosible.com/ontologies) [API reference](https://docs.nosible.com/) [Start trial](https://nosible.com/start-trial) Sources reviewed July 14, 2026 : [[ 1 ] MKT MediaStats product overview](https://www.mktmediastats.com/) ( [dated snapshot](https://archive.is/85YK0); [Wayback copy](https://web.archive.org/web/20260510170725/https://www.mktmediastats.com/) ) , [[ 2 ] MKT MediaStats multi-asset research](https://www.mktmediastats.com/post/multi-asset-rotation) ( [dated snapshot](https://archive.is/EI2sN); [Wayback copy](https://web.archive.org/web/20260416182645/https://www.mktmediastats.com/post/multi-asset-rotation) ) , [[ 3 ] NOSIBLE product overview](https://nosible.com/) ( [dated snapshot](https://archive.is/r7qDv); [Wayback copy](https://web.archive.org/web/20260713185114/https://nosible.com/) ) , [[ 4 ] NOSIBLE SEARCH](https://docs.nosible.com/) ( [dated snapshot](https://archive.is/glqrw) ) , [[ 5 ] NOSIBLE WORLD](https://nosible.world/world) ( [dated snapshot](https://archive.is/tGYc9) ) , [[ 6 ] NOSIBLE WORLD v1.2 trial and coverage](https://nosible.com/start-trial) ( [dated snapshot](https://archive.is/tnkpG) ) , [[ 7 ] NOSIBLE AI-native research overview](https://nosible.com/blog) ( [dated snapshot](https://archive.is/iDZeo) ) , [[ 8 ] NOSIBLE embedding-based research](https://nosible.com/blog/an-embedding-based-approach-to-trade-and-economic-policy-uncertainty) ( [dated snapshot](https://archive.is/1CUvQ) ) , [[ 9 ] NOSIBLE Financial Sentiment v1.2 Base](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) ( [dated snapshot](https://archive.is/7SSz0) ) , [[ 10 ] NOSIBLE Forward-Looking v1.2 Base](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) ( [dated snapshot](https://archive.is/kHxMG) ) . MKT MediaStats, MSCI, and State Street are trademarks of their respective owners. NOSIBLE is not affiliated with, sponsored by, or endorsed by them. Comparison standards, legal context & corrections This comparison was prepared by NOSIBLE, which has a commercial interest in the products being compared. It is based on the cited public materials as they appeared on July 14, 2026. NOSIBLE has not tested every competitor feature, and the page is not a complete statement of either product. Products change; confirm current requirements, availability, and commercial terms with each vendor. Factual claims are attributed to cited first-party materials; evaluative statements reflect NOSIBLE's opinion. Each source includes a dated archive.is snapshot and, where the Internet Archive captured that URL, a timestamp-specific Wayback copy. Archive availability is controlled by those services. Competitor names and marks are used only to identify the products being compared. No affiliation, sponsorship, or endorsement is implied. The FTC says truthful, non-deceptive comparative advertising may identify competitors ( [policy](https://www.ftc.gov/legal-library/browse/statement-policy-regarding-comparative-advertising); [dated snapshot](https://archive.is/nZXY4); [Wayback copy](https://web.archive.org/web/20260215042003/https://www.ftc.gov/legal-library/browse/statement-policy-regarding-comparative-advertising) ). For one U.S. example of nominative-use analysis, see *New Kids on the Block v. News America Publishing, Inc.*, 971 F.2d 302, 308 (9th Cir. 1992) ( [opinion](https://law.justia.com/cases/federal/appellate-courts/F2/971/302/72076/); [dated snapshot](https://archive.is/dnk8j); [Wayback copy](https://web.archive.org/web/20250211213616/https://law.justia.com/cases/federal/appellate-courts/F2/971/302/72076/) ). If you represent MKT MediaStats and believe a factual statement is inaccurate, email [stuart@nosible.com](mailto:stuart@nosible.com?subject=Correction%20request%3A%20NOSIBLE%20vs%20MKT%20MediaStats) with the specific claim and a supporting first-party URL. NOSIBLE will review and correct substantiated errors. > Compare NOSIBLE and MKT MediaStats for point-in-time event intelligence, raw source retrieval, AI agents, factor workflows, delivery, and backtesting. **URL:** https://nosible.com/compare/nosible-vs-mkt-mediastats --- --- title: "NOSIBLE vs Opoint" description: "Compare NOSIBLE and Opoint for media intelligence, curated news data, APIs, dated source evidence, event history, and research-agent workflows." url: "https://nosible.com/compare/nosible-vs-opoint" --- Comparison / Reviewed July 25, 2026 # NOSIBLE vs Opoint Opoint describes a curated global media-intelligence service with portals, feeds, and APIs. [[ 1 a]](https://archive.is/AC7D3) [[ 2 a]](https://archive.is/Pw6R7) NOSIBLE gives agents dated open-web evidence and ranked events when the research record must stay inspectable. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) NOSIBLE-AUTHORED COMPARISON · NOSIBLE HAS A COMMERCIAL INTEREST IN THIS COMPARISON · FACTS ATTRIBUTED TO FIRST-PARTY VENDOR MATERIALS · EVALUATIVE STATEMENTS ARE NOSIBLE'S OPINION · REVIEWED JULY 25, 2026 · DUAL PUBLIC ARCHIVES WHERE SUPPORTED · [STANDARDS & CORRECTIONS](https://nosible.com/compare/nosible-vs-opoint#comparison-standards). IF YOU REPRESENT OPOINT AND BELIEVE A FACTUAL STATEMENT IS INACCURATE, EMAIL [STUART@NOSIBLE.COM](mailto:stuart@nosible.com?subject=Correction%20request%3A%20NOSIBLE%20vs%20Opoint) WITH THE SPECIFIC CLAIM AND A SUPPORTING FIRST-PARTY URL. NOSIBLE WILL REVIEW AND CORRECT SUBSTANTIATED ERRORS. - Opoint says it processes more than three million articles daily and manually curates more than 250,000 sources across 135 languages. [[ 1 a]](https://archive.is/AC7D3) - Opoint publishes delivery through a portal, feeds, and REST/API options. [[ 1 a]](https://archive.is/AC7D3) - Opoint says its global and corporate data is structured and enriched for API use. [[ 1 a]](https://archive.is/AC7D3) [[ 2 a]](https://archive.is/Pw6R7) - NOSIBLE SEARCH and WORLD focus on dated open-web sources, ranked events, and agent workflows, with an embedding per WORLD event for downstream analysis. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) - In NOSIBLE's view, NOSIBLE is the stronger fit when inspectable source timing and event-oriented historical analysis are central. ## A media feed for the newsroom, or evidence for the model? Opoint describes curated global and corporate media data, structured and delivered through a portal, feeds, and APIs. [[ 1 a]](https://archive.is/AC7D3) [[ 2 a]](https://archive.is/Pw6R7) NOSIBLE starts with dated open-web sources and ranked events for research agents. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) In NOSIBLE's view, the decision is less about a generic data count than about the operating model: a broad curated media feed, or inspectable evidence and event context for a research or model workflow. ## Curated media intelligence and inspectable event evidence The competitor column summarizes what the named vendor's cited first-party materials describe; the NOSIBLE column summarizes NOSIBLE's cited materials. Primary use NOSIBLE Dated source and event intelligence [[ 4 a]](https://archive.is/YwGWR) Opoint Curated global and corporate media-intelligence data [[ 2 a]](https://archive.is/Pw6R7) Content scale NOSIBLE Open-web source retrieval; test scope for the task [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 4 a]](https://archive.is/YwGWR) Opoint Opoint says it processes 3M+ articles daily Source curation NOSIBLE Open-web retrieval with source evidence [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Opoint Opoint says it manually curates 250k+ sources Languages NOSIBLE NOSIBLE publishes 95 languages [[ 4 a]](https://archive.is/YwGWR) Opoint Opoint says 135 languages; vendor definitions may differ [[ 1 a]](https://archive.is/AC7D3) Data preparation NOSIBLE Sources, enrichment, and event context Opoint Structured and enriched global and corporate data [[ 2 a]](https://archive.is/Pw6R7) Delivery NOSIBLE API, SDKs, MCP, SEARCH, and WORLD [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) Opoint Portal, feeds, and REST/API options [[ 1 a]](https://archive.is/AC7D3) Timeliness NOSIBLE Dated source and event retrieval [[ 4 a]](https://archive.is/YwGWR) Opoint Opoint says average delivery is under seven minutes [[ 2 a]](https://archive.is/Pw6R7) Event representation NOSIBLE Ranked dated WORLD events [[ 5 a]](https://archive.is/fv0cj) Opoint Media data for downstream intelligence workflows Pricing NOSIBLE Confirm current terms with NOSIBLE Opoint Confirm current terms with Opoint | Dimension | NOSIBLE | Opoint | | --- | --- | --- | | Primary use | Dated source and event intelligence [[ 4 a]](https://archive.is/YwGWR) | Curated global and corporate media-intelligence data [[ 2 a]](https://archive.is/Pw6R7) | | Content scale | Open-web source retrieval; test scope for the task [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 4 a]](https://archive.is/YwGWR) | Opoint says it processes 3M+ articles daily | | Source curation | Open-web retrieval with source evidence [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Opoint says it manually curates 250k+ sources | | Languages | NOSIBLE publishes 95 languages [[ 4 a]](https://archive.is/YwGWR) | Opoint says 135 languages; vendor definitions may differ [[ 1 a]](https://archive.is/AC7D3) | | Data preparation | Sources, enrichment, and event context | Structured and enriched global and corporate data [[ 2 a]](https://archive.is/Pw6R7) | | Delivery | API, SDKs, MCP, SEARCH, and WORLD [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) | Portal, feeds, and REST/API options [[ 1 a]](https://archive.is/AC7D3) | | Timeliness | Dated source and event retrieval [[ 4 a]](https://archive.is/YwGWR) | Opoint says average delivery is under seven minutes [[ 2 a]](https://archive.is/Pw6R7) | | Event representation | Ranked dated WORLD events [[ 5 a]](https://archive.is/fv0cj) | Media data for downstream intelligence workflows | | Pricing | Confirm current terms with NOSIBLE | Confirm current terms with Opoint | ## Monitoring breadth versus evidence depth In NOSIBLE's view, Opoint may be the more direct choice for an organization that requires its curated media corpus, language coverage, or feed and portal delivery. NOSIBLE is built for dated source evidence and ranked events, with an embedding per event for model-oriented analysis. [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) [[ 8 a]](https://archive.is/1CUvQ) In NOSIBLE's view, they can complement one another when media monitoring and agent research are deliberately separate layers. ## Common Opoint comparison questions ### How does NOSIBLE feel it differentiates itself from Opoint? NOSIBLE is an AI-native company with two products: SEARCH and WORLD. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 5 a]](https://archive.is/fv0cj) [[ 7 a]](https://archive.is/iDZeo) SEARCH lets agents find dated open-web sources they can cite and inspect directly. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 4 a]](https://archive.is/YwGWR) WORLD is a live open-web event database for models and backtests, with an embedding per event. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 5 a]](https://archive.is/fv0cj) [[ 6 a]](https://archive.is/tnkpG) [[ 8 a]](https://archive.is/1CUvQ) NOSIBLE is committed to open-source software and makes its models publicly available on Hugging Face. [[ 9 a]](https://archive.is/7SSz0) [[ 10 a]](https://archive.is/kHxMG) Related [WORLD v1.2 trial](https://nosible.com/start-trial#data-coverage) [Sentiment model](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) [Forward-looking model](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) ### What does Opoint publish about its media data? Opoint says it processes more than three million articles daily, manually curates more than 250,000 sources, and covers 135 languages. [[ 1 a]](https://archive.is/AC7D3) It also describes structured and enriched global and corporate data. [[ 2 a]](https://archive.is/Pw6R7) In NOSIBLE's view, Opoint may be the more direct fit when that curated media-data service is the primary requirement. Related [WORLD event database](https://nosible.world/world) ### How do delivery options differ? Opoint publishes portal, feed, and REST/API delivery options for its media intelligence data. [[ 1 a]](https://archive.is/AC7D3) NOSIBLE publishes APIs, SDKs, MCP, SEARCH, and WORLD products. [[ 1 a]](https://archive.is/AC7D3) [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) In NOSIBLE's view, buyers should evaluate integration effort, data contracts, update timing, retention, and permissioning against the workflow that will consume the data. Related [Agentic Search](https://docs.nosible.com/) [WORLD event database](https://nosible.world/world) ### Does Opoint publish a comparable event database? The cited Opoint materials describe curated and enriched media data rather than a published ranked event database comparable to NOSIBLE WORLD. [[ 2 a]](https://archive.is/Pw6R7) [[ 5 a]](https://archive.is/fv0cj) This page therefore does not compare aggregate event counts. In NOSIBLE's view, teams should compare the outputs they actually need, including source records, entities, events, and delivery fields. Related [WORLD event database](https://nosible.world/world) ### How should a team compare language coverage? Opoint says it covers 135 languages, while NOSIBLE publishes a 95-language corpus. [[ 1 a]](https://archive.is/AC7D3) [[ 4 a]](https://archive.is/YwGWR) Counts alone do not establish the relevant source mix or quality. Test the languages, countries, publishers, translation needs, date handling, and retrieval precision that matter for the organization before using either number in a procurement decision. Related [WORLD event database](https://nosible.world/world) ### Can Opoint and NOSIBLE be used together? Potentially. Opoint can provide a curated media-intelligence feed, while NOSIBLE can provide dated open-web sources and ranked event context for agents. [[ 1 a]](https://archive.is/AC7D3) [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) In NOSIBLE's view, teams should document provenance, licensing, duplicates, and timestamps before combining feeds or using them to support research or model features. Related [WORLD event database](https://nosible.world/world) Continue comparing ## Related Vendor Comparisons Compare Opoint with adjacent options for task execution, agent tooling, data access, and research workflows, then evaluate the output shape your application actually needs for production. [Signal AI](https://nosible.com/compare/nosible-vs-signal-ai) [Dow Jones](https://nosible.com/compare/nosible-vs-dow-jones) [AlphaSense](https://nosible.com/compare/nosible-vs-alphasense) [News API](https://nosible.com/compare/nosible-vs-news-api) Dive deeper ## Take your next step today Review the delivered field definitions, classification boundaries, and example values before comparing vendor workflows, so your team can assess what each product actually returns to downstream systems. [Data dictionaries](https://nosible.com/data-dictionaries) [Ontology reference](https://nosible.com/ontologies) [API reference](https://docs.nosible.com/) [Start trial](https://nosible.com/start-trial) Sources reviewed July 25, 2026 : [[ 1 ] Opoint products](https://opoint.com/products) ( [dated snapshot](https://archive.is/AC7D3) ) , [[ 2 ] Opoint data](https://opoint.com/data) ( [dated snapshot](https://archive.is/Pw6R7) ) , [[ 3 ] NOSIBLE product overview](https://nosible.com/) ( [dated snapshot](https://archive.is/8abI7); [Wayback copy](https://web.archive.org/web/20260713185114/https://nosible.com/) ) , [[ 4 ] NOSIBLE SEARCH](https://docs.nosible.com/) ( [dated snapshot](https://archive.is/YwGWR) ) , [[ 5 ] NOSIBLE WORLD](https://nosible.world/world) ( [dated snapshot](https://archive.is/fv0cj) ) , [[ 6 ] NOSIBLE WORLD v1.2 trial and coverage](https://nosible.com/start-trial) ( [dated snapshot](https://archive.is/tnkpG) ) , [[ 7 ] NOSIBLE AI-native research overview](https://nosible.com/blog) ( [dated snapshot](https://archive.is/iDZeo) ) , [[ 8 ] NOSIBLE embedding-based research](https://nosible.com/blog/an-embedding-based-approach-to-trade-and-economic-policy-uncertainty) ( [dated snapshot](https://archive.is/1CUvQ) ) , [[ 9 ] NOSIBLE Financial Sentiment v1.2 Base](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) ( [dated snapshot](https://archive.is/7SSz0) ) , [[ 10 ] NOSIBLE Forward-Looking v1.2 Base](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) ( [dated snapshot](https://archive.is/kHxMG) ) . Opoint is used solely to identify the compared product. NOSIBLE is not affiliated with, sponsored by, or endorsed by Opoint. Comparison standards, legal context & corrections This comparison was prepared by NOSIBLE, which has a commercial interest in the products being compared. It is based on the cited public materials as they appeared on July 25, 2026. NOSIBLE has not tested every competitor feature, and the page is not a complete statement of either product. Products change; confirm current requirements, availability, and commercial terms with each vendor. Factual claims are attributed to cited first-party materials; evaluative statements reflect NOSIBLE's opinion. Each source includes a dated archive.is snapshot and, where the Internet Archive captured that URL, a timestamp-specific Wayback copy. Archive availability is controlled by those services. Competitor names and marks are used only to identify the products being compared. No affiliation, sponsorship, or endorsement is implied. The FTC says truthful, non-deceptive comparative advertising may identify competitors ( [policy](https://www.ftc.gov/legal-library/browse/statement-policy-regarding-comparative-advertising); [dated snapshot](https://archive.is/nZXY4); [Wayback copy](https://web.archive.org/web/20260215042003/https://www.ftc.gov/legal-library/browse/statement-policy-regarding-comparative-advertising) ). For one U.S. example of nominative-use analysis, see *New Kids on the Block v. News America Publishing, Inc.*, 971 F.2d 302, 308 (9th Cir. 1992) ( [opinion](https://law.justia.com/cases/federal/appellate-courts/F2/971/302/72076/); [dated snapshot](https://archive.is/dnk8j); [Wayback copy](https://web.archive.org/web/20250211213616/https://law.justia.com/cases/federal/appellate-courts/F2/971/302/72076/) ). If you represent Opoint and believe a factual statement is inaccurate, email [stuart@nosible.com](mailto:stuart@nosible.com?subject=Correction%20request%3A%20NOSIBLE%20vs%20Opoint) with the specific claim and a supporting first-party URL. NOSIBLE will review and correct substantiated errors. > Compare NOSIBLE and Opoint for media intelligence, curated news data, APIs, dated source evidence, event history, and research-agent workflows. **URL:** https://nosible.com/compare/nosible-vs-opoint --- --- title: "NOSIBLE vs Alexandria" description: "Compare NOSIBLE and Alexandria Technology across AI agent readiness, point-in-time retrieval, sentiment scoring, event data, APIs, and source coverage." url: "https://nosible.com/compare/nosible-vs-alexandria" --- Comparison / Reviewed July 14, 2026 # NOSIBLE vs Alexandria Alexandria applies financial NLP to news, earnings calls, social media, macro news, and ESG content. [[ 1 a]](https://archive.is/5FgEz) [[ 1 b]](https://web.archive.org/web/20260510103526/https://www.alexandriatechnology.com/) NOSIBLE gives agents a different dated source-and-event layer. [[ 4 a]](https://archive.is/glqrw) NOSIBLE-AUTHORED COMPARISON · NOSIBLE HAS A COMMERCIAL INTEREST IN THIS COMPARISON · FACTS ATTRIBUTED TO FIRST-PARTY VENDOR MATERIALS · EVALUATIVE STATEMENTS ARE NOSIBLE'S OPINION · REVIEWED JULY 14, 2026 · DUAL PUBLIC ARCHIVES WHERE SUPPORTED · [STANDARDS & CORRECTIONS](https://nosible.com/compare/nosible-vs-alexandria#comparison-standards). IF YOU REPRESENT ALEXANDRIA TECHNOLOGY AND BELIEVE A FACTUAL STATEMENT IS INACCURATE, EMAIL [STUART@NOSIBLE.COM](mailto:stuart@nosible.com?subject=Correction%20request%3A%20NOSIBLE%20vs%20Alexandria) WITH THE SPECIFIC CLAIM AND A SUPPORTING FIRST-PARTY URL. NOSIBLE WILL REVIEW AND CORRECT SUBSTANTIATED ERRORS. - Alexandria says its NLP identifies entities, topics, and sentiment in financial text. [[ 1 a]](https://archive.is/5FgEz) [[ 1 b]](https://web.archive.org/web/20260510103526/https://www.alexandriatechnology.com/) - Its published product areas include earnings calls, company news, social media, macro news, and ESG. [[ 2 a]](https://archive.is/e5H3r) [[ 2 b]](https://web.archive.org/web/20260510111357/https://www.alexandriatechnology.com/company-news) - Alexandria's Company News page describes millions of premium-publisher articles and more than 20 years of historical data. [[ 2 a]](https://archive.is/e5H3r) [[ 2 b]](https://web.archive.org/web/20260510111357/https://www.alexandriatechnology.com/company-news) - NOSIBLE emphasizes dated source retrieval, ranked events, and agent workflows across the open web. [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) - In NOSIBLE's view, consider NOSIBLE when the system needs inspectable evidence as well as enrichment. ## Financial NLP and source-level evidence Alexandria is a specialist financial-NLP provider with sentiment and classification products for investment professionals. [[ 1 a]](https://archive.is/5FgEz) [[ 1 b]](https://web.archive.org/web/20260510103526/https://www.alexandriatechnology.com/) NOSIBLE lets agents and research systems search dated open-web material, inspect underlying text, and retrieve ranked events. [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) In NOSIBLE's view, NOSIBLE is the stronger fit when auditability and source retrieval are as important as the derived score. ## Financial NLP and source-level evidence The competitor column summarizes what the named vendor's cited first-party materials describe; the NOSIBLE column summarizes NOSIBLE's cited materials. Primary use NOSIBLE Search and event intelligence [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Alexandria Technology Financial NLP scoring [[ 1 a]](https://archive.is/5FgEz) [[ 1 b]](https://web.archive.org/web/20260510103526/https://www.alexandriatechnology.com/) Illustrative users NOSIBLE AI agents, backtests, research systems [[ 4 a]](https://archive.is/glqrw) Alexandria Technology Investment teams using text analytics [[ 1 a]](https://archive.is/5FgEz) [[ 1 b]](https://web.archive.org/web/20260510103526/https://www.alexandriatechnology.com/) Primary output NOSIBLE Dated documents and ranked events [[ 5 a]](https://archive.is/tGYc9) Alexandria Technology Entity, topic, and sentiment analytics [[ 1 a]](https://archive.is/5FgEz) [[ 1 b]](https://web.archive.org/web/20260510103526/https://www.alexandriatechnology.com/) Content scope NOSIBLE News, corporate, government text [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Alexandria Technology Earnings calls, company and macro news, social media, and ESG content [[ 1 a]](https://archive.is/5FgEz) [[ 1 b]](https://web.archive.org/web/20260510103526/https://www.alexandriatechnology.com/) Company mapping and coverage NOSIBLE Ticker and organization mappings linked to source evidence [[ 5 a]](https://archive.is/tGYc9) Alexandria Technology Alexandria says it covers every publicly traded company globally [[ 1 a]](https://archive.is/5FgEz) [[ 1 b]](https://web.archive.org/web/20260510103526/https://www.alexandriatechnology.com/) History NOSIBLE Roughly 30 years [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Alexandria Technology 20+ years of historical data [[ 2 a]](https://archive.is/e5H3r) [[ 2 b]](https://web.archive.org/web/20260510111357/https://www.alexandriatechnology.com/company-news) Point-in-time method NOSIBLE Five-way date verification [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Alexandria Technology Historical and real-time data [[ 1 a]](https://archive.is/5FgEz) [[ 1 b]](https://web.archive.org/web/20260510103526/https://www.alexandriatechnology.com/) Delivery NOSIBLE API, SDKs, MCP, agentic search [[ 4 a]](https://archive.is/glqrw) Alexandria Technology Data partnerships and investment-workflow delivery [[ 1 a]](https://archive.is/5FgEz) [[ 1 b]](https://web.archive.org/web/20260510103526/https://www.alexandriatechnology.com/) Agent readiness NOSIBLE Built for AI agents [[ 4 a]](https://archive.is/glqrw) Alexandria Technology Built for investment professionals and quantitative workflows [[ 1 a]](https://archive.is/5FgEz) [[ 1 b]](https://web.archive.org/web/20260510103526/https://www.alexandriatechnology.com/) | Dimension | NOSIBLE | Alexandria Technology | | --- | --- | --- | | Primary use | Search and event intelligence [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Financial NLP scoring [[ 1 a]](https://archive.is/5FgEz) [[ 1 b]](https://web.archive.org/web/20260510103526/https://www.alexandriatechnology.com/) | | Illustrative users | AI agents, backtests, research systems [[ 4 a]](https://archive.is/glqrw) | Investment teams using text analytics [[ 1 a]](https://archive.is/5FgEz) [[ 1 b]](https://web.archive.org/web/20260510103526/https://www.alexandriatechnology.com/) | | Primary output | Dated documents and ranked events [[ 5 a]](https://archive.is/tGYc9) | Entity, topic, and sentiment analytics [[ 1 a]](https://archive.is/5FgEz) [[ 1 b]](https://web.archive.org/web/20260510103526/https://www.alexandriatechnology.com/) | | Content scope | News, corporate, government text [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Earnings calls, company and macro news, social media, and ESG content [[ 1 a]](https://archive.is/5FgEz) [[ 1 b]](https://web.archive.org/web/20260510103526/https://www.alexandriatechnology.com/) | | Company mapping and coverage | Ticker and organization mappings linked to source evidence [[ 5 a]](https://archive.is/tGYc9) | Alexandria says it covers every publicly traded company globally [[ 1 a]](https://archive.is/5FgEz) [[ 1 b]](https://web.archive.org/web/20260510103526/https://www.alexandriatechnology.com/) | | History | Roughly 30 years [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | 20+ years of historical data [[ 2 a]](https://archive.is/e5H3r) [[ 2 b]](https://web.archive.org/web/20260510111357/https://www.alexandriatechnology.com/company-news) | | Point-in-time method | Five-way date verification [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Historical and real-time data [[ 1 a]](https://archive.is/5FgEz) [[ 1 b]](https://web.archive.org/web/20260510103526/https://www.alexandriatechnology.com/) | | Delivery | API, SDKs, MCP, agentic search [[ 4 a]](https://archive.is/glqrw) | Data partnerships and investment-workflow delivery [[ 1 a]](https://archive.is/5FgEz) [[ 1 b]](https://web.archive.org/web/20260510103526/https://www.alexandriatechnology.com/) | | Agent readiness | Built for AI agents [[ 4 a]](https://archive.is/glqrw) | Built for investment professionals and quantitative workflows [[ 1 a]](https://archive.is/5FgEz) [[ 1 b]](https://web.archive.org/web/20260510103526/https://www.alexandriatechnology.com/) | ## Potential fit by workflow Alexandria's cited materials describe specialist financial-NLP analytics over supported content products. [[ 1 a]](https://archive.is/5FgEz) [[ 1 b]](https://web.archive.org/web/20260510103526/https://www.alexandriatechnology.com/) NOSIBLE's cited materials describe cross-domain dated documents, event retrieval, and source-level context for agents. [[ 4 a]](https://archive.is/glqrw) In NOSIBLE's view, consider the product whose published output matches the requirement; a team may also use NOSIBLE evidence alongside a specialist classifier. ## Common Alexandria comparison questions ### How does NOSIBLE feel it differentiates itself from Alexandria Technology? NOSIBLE is an AI-native company with two products: SEARCH and WORLD. [[ 5 a]](https://archive.is/tGYc9) [[ 7 a]](https://archive.is/iDZeo) SEARCH lets agents find dated open-web sources they can cite and inspect directly. [[ 4 a]](https://archive.is/glqrw) WORLD is a live open-web event database for models and backtests, with an embedding per event. [[ 5 a]](https://archive.is/tGYc9) [[ 6 a]](https://archive.is/tnkpG) [[ 8 a]](https://archive.is/1CUvQ) NOSIBLE is committed to open-source software and makes its models publicly available on Hugging Face. [[ 9 a]](https://archive.is/7SSz0) [[ 10 a]](https://archive.is/kHxMG) Related [WORLD v1.2 trial](https://nosible.com/start-trial#data-coverage) [Sentiment model](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) [Forward-looking model](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) ### Are we buying document scores or the underlying source layer? Alexandria publishes financial-NLP outputs; NOSIBLE publishes dated source-retrieval, ranked-event, entity-context, and agent workflows. [[ 1 a]](https://archive.is/5FgEz) [[ 1 b]](https://web.archive.org/web/20260510103526/https://www.alexandriatechnology.com/) [[ 4 a]](https://archive.is/glqrw) In NOSIBLE's view, Alexandria is the more direct option for the first requirement, while NOSIBLE is the better fit for the second. Buyers should test both on the same document sample. Related [WORLD event database](https://nosible.world/world) ### What if our workflow already uses FactSet earnings-call data? Alexandria's homepage says its Earnings Calls product covers global corporate events from FactSet. [[ 1 a]](https://archive.is/5FgEz) [[ 1 b]](https://web.archive.org/web/20260510103526/https://www.alexandriatechnology.com/) NOSIBLE maps open-web events to tickers and other entities while retaining source context. [[ 5 a]](https://archive.is/tGYc9) In NOSIBLE's view, buyers should test identifier compatibility, source coverage, and access to the original evidence required by the workflow. Related [WORLD event database](https://nosible.world/world) ### Can NOSIBLE replace Alexandria Transcript Text Analytics? NOSIBLE is not a transcript-feed replacement. [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) In NOSIBLE's view, Alexandria should remain under consideration when earnings-call analytics are the requirement. NOSIBLE can provide surrounding dated company, government, regional, and specialist-source context for a cross-domain event workflow. [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Related [WORLD event database](https://nosible.world/world) ### How do Alexandria's analyst-trained classifiers compare with NOSIBLE for auditability? Ask for the text identifier, timestamp, entity mapping, model output, and revision policy behind each score. NOSIBLE emphasizes retrievable source evidence and replayable events. [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Alexandria emphasizes specialist financial classification. [[ 1 a]](https://archive.is/5FgEz) [[ 1 b]](https://web.archive.org/web/20260510103526/https://www.alexandriatechnology.com/) Auditability should be tested at the record level in both products. Related [Sentiment model](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) [Sentiment research](https://nosible.com/blog/fast-enough-to-matter-productionizing-tiny-transformers-for-signal-extraction) [WORLD event database](https://nosible.world/world) ### What evidence should we ask for before using either product in backtests? Ask whether the system can replay the source material and derived fields available at a simulated date, and how later corrections are handled. Run a fixed set of historical queries and record any revisions. Product descriptions alone do not establish backtest suitability. Related [WORLD event database](https://nosible.world/world) Continue comparing ## Related Vendor Comparisons Compare Alexandria Technology with adjacent options for task execution, agent tooling, data access, and research workflows, then evaluate the output shape your application actually needs for production. [SESAMm](https://nosible.com/compare/nosible-vs-sesamm) [MKT MediaStats](https://nosible.com/compare/nosible-vs-mkt-mediastats) [Signal AI](https://nosible.com/compare/nosible-vs-signal-ai) [MarketPsych](https://nosible.com/compare/nosible-vs-marketpsych) Dive deeper ## Take your next step today Review the delivered field definitions, classification boundaries, and example values before comparing vendor workflows, so your team can assess what each product actually returns to downstream systems. [Data dictionaries](https://nosible.com/data-dictionaries) [Ontology reference](https://nosible.com/ontologies) [API reference](https://docs.nosible.com/) [Start trial](https://nosible.com/start-trial) Sources reviewed July 14, 2026 : [[ 1 ] Alexandria Technology product overview](https://www.alexandriatechnology.com/) ( [dated snapshot](https://archive.is/5FgEz); [Wayback copy](https://web.archive.org/web/20260510103526/https://www.alexandriatechnology.com/) ) , [[ 2 ] Alexandria Company News](https://www.alexandriatechnology.com/company-news) ( [dated snapshot](https://archive.is/e5H3r); [Wayback copy](https://web.archive.org/web/20260510111357/https://www.alexandriatechnology.com/company-news) ) , [[ 3 ] NOSIBLE product overview](https://nosible.com/) ( [dated snapshot](https://archive.is/r7qDv); [Wayback copy](https://web.archive.org/web/20260713185114/https://nosible.com/) ) , [[ 4 ] NOSIBLE SEARCH](https://docs.nosible.com/) ( [dated snapshot](https://archive.is/glqrw) ) , [[ 5 ] NOSIBLE WORLD](https://nosible.world/world) ( [dated snapshot](https://archive.is/tGYc9) ) , [[ 6 ] NOSIBLE WORLD v1.2 trial and coverage](https://nosible.com/start-trial) ( [dated snapshot](https://archive.is/tnkpG) ) , [[ 7 ] NOSIBLE AI-native research overview](https://nosible.com/blog) ( [dated snapshot](https://archive.is/iDZeo) ) , [[ 8 ] NOSIBLE embedding-based research](https://nosible.com/blog/an-embedding-based-approach-to-trade-and-economic-policy-uncertainty) ( [dated snapshot](https://archive.is/1CUvQ) ) , [[ 9 ] NOSIBLE Financial Sentiment v1.2 Base](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) ( [dated snapshot](https://archive.is/7SSz0) ) , [[ 10 ] NOSIBLE Forward-Looking v1.2 Base](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) ( [dated snapshot](https://archive.is/kHxMG) ) . Alexandria Technology and FactSet are trademarks of their respective owners. NOSIBLE is not affiliated with, sponsored by, or endorsed by Alexandria Technology or FactSet. Comparison standards, legal context & corrections This comparison was prepared by NOSIBLE, which has a commercial interest in the products being compared. It is based on the cited public materials as they appeared on July 14, 2026. NOSIBLE has not tested every competitor feature, and the page is not a complete statement of either product. Products change; confirm current requirements, availability, and commercial terms with each vendor. Factual claims are attributed to cited first-party materials; evaluative statements reflect NOSIBLE's opinion. Each source includes a dated archive.is snapshot and, where the Internet Archive captured that URL, a timestamp-specific Wayback copy. Archive availability is controlled by those services. Competitor names and marks are used only to identify the products being compared. No affiliation, sponsorship, or endorsement is implied. The FTC says truthful, non-deceptive comparative advertising may identify competitors ( [policy](https://www.ftc.gov/legal-library/browse/statement-policy-regarding-comparative-advertising); [dated snapshot](https://archive.is/nZXY4); [Wayback copy](https://web.archive.org/web/20260215042003/https://www.ftc.gov/legal-library/browse/statement-policy-regarding-comparative-advertising) ). For one U.S. example of nominative-use analysis, see *New Kids on the Block v. News America Publishing, Inc.*, 971 F.2d 302, 308 (9th Cir. 1992) ( [opinion](https://law.justia.com/cases/federal/appellate-courts/F2/971/302/72076/); [dated snapshot](https://archive.is/dnk8j); [Wayback copy](https://web.archive.org/web/20250211213616/https://law.justia.com/cases/federal/appellate-courts/F2/971/302/72076/) ). If you represent Alexandria Technology and believe a factual statement is inaccurate, email [stuart@nosible.com](mailto:stuart@nosible.com?subject=Correction%20request%3A%20NOSIBLE%20vs%20Alexandria) with the specific claim and a supporting first-party URL. NOSIBLE will review and correct substantiated errors. > Compare NOSIBLE and Alexandria Technology across AI agent readiness, point-in-time retrieval, sentiment scoring, event data, APIs, and source coverage. **URL:** https://nosible.com/compare/nosible-vs-alexandria --- --- title: "NOSIBLE vs Exa" description: "Compare NOSIBLE and Exa for web search APIs, agent workflows, structured web-data collection, source evidence, event history, and research delivery." url: "https://nosible.com/compare/nosible-vs-exa" --- Comparison / Reviewed July 25, 2026 # NOSIBLE vs Exa Exa documents web search and Websets for building structured collections from the web. [[ 1 a]](https://archive.is/VOQpn) [[ 1 b]](https://web.archive.org/web/20260725105148/https://exa.ai/docs/reference/search) [[ 2 a]](https://archive.is/nxuP7) [[ 2 b]](https://web.archive.org/web/20260725105209/https://exa.ai/docs/websets/api-guide) NOSIBLE pairs dated source retrieval with ranked events when the output needs to explain a change over time. [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) NOSIBLE-AUTHORED COMPARISON · NOSIBLE HAS A COMMERCIAL INTEREST IN THIS COMPARISON · FACTS ATTRIBUTED TO FIRST-PARTY VENDOR MATERIALS · EVALUATIVE STATEMENTS ARE NOSIBLE'S OPINION · REVIEWED JULY 25, 2026 · DUAL PUBLIC ARCHIVES WHERE SUPPORTED · [STANDARDS & CORRECTIONS](https://nosible.com/compare/nosible-vs-exa#comparison-standards). IF YOU REPRESENT EXA AND BELIEVE A FACTUAL STATEMENT IS INACCURATE, EMAIL [STUART@NOSIBLE.COM](mailto:stuart@nosible.com?subject=Correction%20request%3A%20NOSIBLE%20vs%20Exa) WITH THE SPECIFIC CLAIM AND A SUPPORTING FIRST-PARTY URL. NOSIBLE WILL REVIEW AND CORRECT SUBSTANTIATED ERRORS. - Exa documents a Search API that searches the web and can return extracted result content. [[ 1 a]](https://archive.is/VOQpn) [[ 1 b]](https://web.archive.org/web/20260725105148/https://exa.ai/docs/reference/search) - Exa says Websets organize discovered web content into structured items that can be verified and enriched. [[ 2 a]](https://archive.is/nxuP7) [[ 2 b]](https://web.archive.org/web/20260725105209/https://exa.ai/docs/websets/api-guide) - NOSIBLE SEARCH emphasizes dated open-web evidence, while WORLD supplies ranked events with an embedding per event for model and backtest workflows. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 5 a]](https://archive.is/fv0cj) [[ 8 a]](https://archive.is/1CUvQ) - In NOSIBLE's view, the products overlap for web retrieval but publish different downstream data models. - In NOSIBLE's view, NOSIBLE is the stronger fit when event history and replayable source timing are part of the required output. ## Build a collection, or investigate a change? Exa documents web search and Websets: search agents discover content, then verification and enrichment turn it into structured items. [[ 1 a]](https://archive.is/VOQpn) [[ 1 b]](https://web.archive.org/web/20260725105148/https://exa.ai/docs/reference/search) [[ 2 a]](https://archive.is/nxuP7) [[ 2 b]](https://web.archive.org/web/20260725105209/https://exa.ai/docs/websets/api-guide) NOSIBLE takes a different shape—dated sources in SEARCH and a separate ranked event layer in WORLD. [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) In NOSIBLE's view, a buyer should start with the output they need: a custom collection, or an evidence-backed account of what changed and when. ## Custom web collections and time-bound event research The competitor column summarizes what the named vendor's cited first-party materials describe; the NOSIBLE column summarizes NOSIBLE's cited materials. Primary use NOSIBLE Dated source and event intelligence [[ 4 a]](https://archive.is/YwGWR) Exa Programmatic web search and structured web-data collection [[ 1 a]](https://archive.is/VOQpn) [[ 1 b]](https://web.archive.org/web/20260725105148/https://exa.ai/docs/reference/search) [[ 2 a]](https://archive.is/nxuP7) [[ 2 b]](https://web.archive.org/web/20260725105209/https://exa.ai/docs/websets/api-guide) Search output NOSIBLE Sources with dated evidence Exa Search results with optional extracted content [[ 1 a]](https://archive.is/VOQpn) [[ 1 b]](https://web.archive.org/web/20260725105148/https://exa.ai/docs/reference/search) Collection model NOSIBLE Queries, sources, and ranked events [[ 5 a]](https://archive.is/fv0cj) Exa Websets, items, searches, and enrichments [[ 2 a]](https://archive.is/nxuP7) [[ 2 b]](https://web.archive.org/web/20260725105209/https://exa.ai/docs/websets/api-guide) Verification NOSIBLE Documented buyer evaluation Exa Exa describes Webset item verification against criteria [[ 2 a]](https://archive.is/nxuP7) [[ 2 b]](https://web.archive.org/web/20260725105209/https://exa.ai/docs/websets/api-guide) Entity workflow NOSIBLE WORLD entity and ticker context [[ 5 a]](https://archive.is/fv0cj) Exa Structured item fields and type-specific enrichments [[ 2 a]](https://archive.is/nxuP7) [[ 2 b]](https://web.archive.org/web/20260725105209/https://exa.ai/docs/websets/api-guide) Point-in-time method NOSIBLE Replayable source and event history [[ 5 a]](https://archive.is/fv0cj) Exa Evaluate historical behavior with a fixed test Event representation NOSIBLE Ranked dated WORLD events [[ 5 a]](https://archive.is/fv0cj) Exa Web content collected for a user-defined Webset Delivery NOSIBLE API, SDKs, MCP, SEARCH, and WORLD [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) Exa Search and Websets APIs [[ 2 a]](https://archive.is/nxuP7) [[ 2 b]](https://web.archive.org/web/20260725105209/https://exa.ai/docs/websets/api-guide) Pricing NOSIBLE Confirm current terms with NOSIBLE Exa Confirm current terms with Exa | Dimension | NOSIBLE | Exa | | --- | --- | --- | | Primary use | Dated source and event intelligence [[ 4 a]](https://archive.is/YwGWR) | Programmatic web search and structured web-data collection [[ 1 a]](https://archive.is/VOQpn) [[ 1 b]](https://web.archive.org/web/20260725105148/https://exa.ai/docs/reference/search) [[ 2 a]](https://archive.is/nxuP7) [[ 2 b]](https://web.archive.org/web/20260725105209/https://exa.ai/docs/websets/api-guide) | | Search output | Sources with dated evidence | Search results with optional extracted content [[ 1 a]](https://archive.is/VOQpn) [[ 1 b]](https://web.archive.org/web/20260725105148/https://exa.ai/docs/reference/search) | | Collection model | Queries, sources, and ranked events [[ 5 a]](https://archive.is/fv0cj) | Websets, items, searches, and enrichments [[ 2 a]](https://archive.is/nxuP7) [[ 2 b]](https://web.archive.org/web/20260725105209/https://exa.ai/docs/websets/api-guide) | | Verification | Documented buyer evaluation | Exa describes Webset item verification against criteria [[ 2 a]](https://archive.is/nxuP7) [[ 2 b]](https://web.archive.org/web/20260725105209/https://exa.ai/docs/websets/api-guide) | | Entity workflow | WORLD entity and ticker context [[ 5 a]](https://archive.is/fv0cj) | Structured item fields and type-specific enrichments [[ 2 a]](https://archive.is/nxuP7) [[ 2 b]](https://web.archive.org/web/20260725105209/https://exa.ai/docs/websets/api-guide) | | Point-in-time method | Replayable source and event history [[ 5 a]](https://archive.is/fv0cj) | Evaluate historical behavior with a fixed test | | Event representation | Ranked dated WORLD events [[ 5 a]](https://archive.is/fv0cj) | Web content collected for a user-defined Webset | | Delivery | API, SDKs, MCP, SEARCH, and WORLD [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) | Search and Websets APIs [[ 2 a]](https://archive.is/nxuP7) [[ 2 b]](https://web.archive.org/web/20260725105209/https://exa.ai/docs/websets/api-guide) | | Pricing | Confirm current terms with NOSIBLE | Confirm current terms with Exa | ## Collection building versus time-bound investigation In NOSIBLE's view, Exa may be the more direct choice when a team wants to discover entities, verify them against criteria, and enrich a purpose-built Webset. NOSIBLE is designed around dated source evidence and ranked events, including an embedding per event for downstream models. [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) [[ 8 a]](https://archive.is/1CUvQ) In NOSIBLE's view, the two can fit together when collection building informs event-oriented historical analysis. ## Common Exa comparison questions ### How does NOSIBLE feel it differentiates itself from Exa? NOSIBLE is an AI-native company with two products: SEARCH and WORLD. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 5 a]](https://archive.is/fv0cj) [[ 7 a]](https://archive.is/iDZeo) SEARCH lets agents find dated open-web sources they can cite and inspect directly. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 4 a]](https://archive.is/YwGWR) WORLD is a live open-web event database for models and backtests, with an embedding per event. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 5 a]](https://archive.is/fv0cj) [[ 6 a]](https://archive.is/tnkpG) [[ 8 a]](https://archive.is/1CUvQ) NOSIBLE is committed to open-source software and makes its models publicly available on Hugging Face. [[ 9 a]](https://archive.is/7SSz0) [[ 10 a]](https://archive.is/kHxMG) Related [WORLD v1.2 trial](https://nosible.com/start-trial#data-coverage) [Sentiment model](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) [Forward-looking model](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) ### How does NOSIBLE differ from Exa Search? Exa documents a Search API that searches the web and can return extracted result content. [[ 1 a]](https://archive.is/VOQpn) [[ 1 b]](https://web.archive.org/web/20260725105148/https://exa.ai/docs/reference/search) NOSIBLE SEARCH emphasizes dated source retrieval for agents and research systems. [[ 4 a]](https://archive.is/YwGWR) In NOSIBLE's view, the practical comparison is not only search quality: buyers should test source timing, inspection needs, result structure, and the downstream workflow each API supports. Related [Agentic Search](https://docs.nosible.com/) ### What is Exa Websets designed to do? Exa says Websets organize discovered web content into structured items, then support verification and enrichment against user-defined criteria. [[ 2 a]](https://archive.is/nxuP7) [[ 2 b]](https://web.archive.org/web/20260725105209/https://exa.ai/docs/websets/api-guide) NOSIBLE WORLD is a separate ranked event database rather than a general-purpose Webset container. [[ 5 a]](https://archive.is/fv0cj) In NOSIBLE's view, Websets may be the more direct fit for building a custom researched entity collection. Related [WORLD event database](https://nosible.world/world) ### Can Exa and NOSIBLE be complementary? Potentially. Exa publishes search, collection, verification, and enrichment workflows for web data. [[ 2 a]](https://archive.is/nxuP7) [[ 2 b]](https://web.archive.org/web/20260725105209/https://exa.ai/docs/websets/api-guide) NOSIBLE publishes dated source retrieval and ranked event history. [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) In NOSIBLE's view, a team could use Exa for a custom entity-collection process while using NOSIBLE where an event layer, source timing, or historical replay is central. Related [WORLD event database](https://nosible.world/world) ### How should a team evaluate point-in-time behavior? Run identical historical questions through both products and record the returned URLs, publication dates, extraction timing, and available structured fields. Exa documents search and Webset objects, while NOSIBLE emphasizes dated sources and events. [[ 4 a]](https://archive.is/YwGWR) A reproducible test is more useful than inferring historical behavior from a current product description alone. Related [WORLD event database](https://nosible.world/world) ### Does Exa publish comparable entity or event counts? The cited Exa materials describe Websets, items, searches, verification, and enrichments, not a directly comparable published total for NOSIBLE WORLD entities or events. [[ 2 a]](https://archive.is/nxuP7) [[ 2 b]](https://web.archive.org/web/20260725105209/https://exa.ai/docs/websets/api-guide) [[ 5 a]](https://archive.is/fv0cj) This page therefore does not compare aggregate entity counts. In NOSIBLE's view, buyers should compare definitions and task-level results before treating collection sizes as equivalent. Related [WORLD event database](https://nosible.world/world) Continue comparing ## Related Vendor Comparisons Compare Exa with adjacent options for task execution, agent tooling, data access, and research workflows, then evaluate the output shape your application actually needs for production. [Parallel](https://nosible.com/compare/nosible-vs-parallel) [Tavily](https://nosible.com/compare/nosible-vs-tavily) [Common Crawl](https://nosible.com/compare/nosible-vs-common-crawl) [News API](https://nosible.com/compare/nosible-vs-news-api) Dive deeper ## Take your next step today Review the delivered field definitions, classification boundaries, and example values before comparing vendor workflows, so your team can assess what each product actually returns to downstream systems. [Data dictionaries](https://nosible.com/data-dictionaries) [Ontology reference](https://nosible.com/ontologies) [API reference](https://docs.nosible.com/) [Start trial](https://nosible.com/start-trial) Sources reviewed July 25, 2026 : [[ 1 ] Exa Search API](https://exa.ai/docs/reference/search) ( [dated snapshot](https://archive.is/VOQpn); [Wayback copy](https://web.archive.org/web/20260725105148/https://exa.ai/docs/reference/search) ) , [[ 2 ] Exa Websets API](https://exa.ai/docs/websets/api-guide) ( [dated snapshot](https://archive.is/nxuP7); [Wayback copy](https://web.archive.org/web/20260725105209/https://exa.ai/docs/websets/api-guide) ) , [[ 3 ] NOSIBLE product overview](https://nosible.com/) ( [dated snapshot](https://archive.is/8abI7); [Wayback copy](https://web.archive.org/web/20260713185114/https://nosible.com/) ) , [[ 4 ] NOSIBLE SEARCH](https://docs.nosible.com/) ( [dated snapshot](https://archive.is/YwGWR) ) , [[ 5 ] NOSIBLE WORLD](https://nosible.world/world) ( [dated snapshot](https://archive.is/fv0cj) ) , [[ 6 ] NOSIBLE WORLD v1.2 trial and coverage](https://nosible.com/start-trial) ( [dated snapshot](https://archive.is/tnkpG) ) , [[ 7 ] NOSIBLE AI-native research overview](https://nosible.com/blog) ( [dated snapshot](https://archive.is/iDZeo) ) , [[ 8 ] NOSIBLE embedding-based research](https://nosible.com/blog/an-embedding-based-approach-to-trade-and-economic-policy-uncertainty) ( [dated snapshot](https://archive.is/1CUvQ) ) , [[ 9 ] NOSIBLE Financial Sentiment v1.2 Base](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) ( [dated snapshot](https://archive.is/7SSz0) ) , [[ 10 ] NOSIBLE Forward-Looking v1.2 Base](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) ( [dated snapshot](https://archive.is/kHxMG) ) . Exa is used solely to identify the compared product. NOSIBLE is not affiliated with, sponsored by, or endorsed by Exa. Comparison standards, legal context & corrections This comparison was prepared by NOSIBLE, which has a commercial interest in the products being compared. It is based on the cited public materials as they appeared on July 25, 2026. NOSIBLE has not tested every competitor feature, and the page is not a complete statement of either product. Products change; confirm current requirements, availability, and commercial terms with each vendor. Factual claims are attributed to cited first-party materials; evaluative statements reflect NOSIBLE's opinion. Each source includes a dated archive.is snapshot and, where the Internet Archive captured that URL, a timestamp-specific Wayback copy. Archive availability is controlled by those services. Competitor names and marks are used only to identify the products being compared. No affiliation, sponsorship, or endorsement is implied. The FTC says truthful, non-deceptive comparative advertising may identify competitors ( [policy](https://www.ftc.gov/legal-library/browse/statement-policy-regarding-comparative-advertising); [dated snapshot](https://archive.is/nZXY4); [Wayback copy](https://web.archive.org/web/20260215042003/https://www.ftc.gov/legal-library/browse/statement-policy-regarding-comparative-advertising) ). For one U.S. example of nominative-use analysis, see *New Kids on the Block v. News America Publishing, Inc.*, 971 F.2d 302, 308 (9th Cir. 1992) ( [opinion](https://law.justia.com/cases/federal/appellate-courts/F2/971/302/72076/); [dated snapshot](https://archive.is/dnk8j); [Wayback copy](https://web.archive.org/web/20250211213616/https://law.justia.com/cases/federal/appellate-courts/F2/971/302/72076/) ). If you represent Exa and believe a factual statement is inaccurate, email [stuart@nosible.com](mailto:stuart@nosible.com?subject=Correction%20request%3A%20NOSIBLE%20vs%20Exa) with the specific claim and a supporting first-party URL. NOSIBLE will review and correct substantiated errors. > Compare NOSIBLE and Exa for web search APIs, agent workflows, structured web-data collection, source evidence, event history, and research delivery. **URL:** https://nosible.com/compare/nosible-vs-exa --- --- title: "NOSIBLE vs Parallel Web Systems" description: "Compare NOSIBLE and Parallel Web Systems for web research APIs, task execution, source evidence, citations, event history, and agent workflows." url: "https://nosible.com/compare/nosible-vs-parallel" --- Comparison / Reviewed July 25, 2026 # NOSIBLE vs Parallel Web Systems Parallel Web Systems documents APIs that execute web-research tasks. NOSIBLE supplies the dated source and ranked-event layer for teams that need to inspect the evidence behind a research workflow. [[ 4 a]](https://archive.is/YwGWR) NOSIBLE-AUTHORED COMPARISON · NOSIBLE HAS A COMMERCIAL INTEREST IN THIS COMPARISON · FACTS ATTRIBUTED TO FIRST-PARTY VENDOR MATERIALS · EVALUATIVE STATEMENTS ARE NOSIBLE'S OPINION · REVIEWED JULY 25, 2026 · DUAL PUBLIC ARCHIVES WHERE SUPPORTED · [STANDARDS & CORRECTIONS](https://nosible.com/compare/nosible-vs-parallel#comparison-standards). IF YOU REPRESENT PARALLEL WEB SYSTEMS AND BELIEVE A FACTUAL STATEMENT IS INACCURATE, EMAIL [STUART@NOSIBLE.COM](mailto:stuart@nosible.com?subject=Correction%20request%3A%20NOSIBLE%20vs%20Parallel%20Web%20Systems) WITH THE SPECIFIC CLAIM AND A SUPPORTING FIRST-PARTY URL. NOSIBLE WILL REVIEW AND CORRECT SUBSTANTIATED ERRORS. - Parallel documents a Search API that takes a natural-language objective and returns LLM-optimized excerpts. [[ 1 a]](https://archive.is/cN9EP) [[ 1 b]](https://web.archive.org/web/20260725105649/https://docs.parallel.ai/search/search-quickstart) - Parallel says its Task API combines AI inference, web search, live crawling, and structured output with citations. [[ 2 a]](https://archive.is/Rltif) [[ 2 b]](https://web.archive.org/web/20260725105732/https://docs.parallel.ai/task-api/task-quickstart) - NOSIBLE SEARCH and WORLD keep dated open-web sources and ranked event history available for downstream agents rather than delegating the full task to one endpoint. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) - In NOSIBLE's view, Parallel may be the more direct fit when an application needs a cited research task completed through one API call. - In NOSIBLE's view, NOSIBLE is the stronger fit when replayable event history and inspected source timing are required outputs. ## A finished research task or a durable evidence layer Parallel documents APIs that turn a natural-language objective into search results or a structured, cited research task. [[ 1 a]](https://archive.is/cN9EP) [[ 1 b]](https://web.archive.org/web/20260725105649/https://docs.parallel.ai/search/search-quickstart) NOSIBLE supplies dated sources and ranked events for a team to use inside its own research system. [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) In NOSIBLE's view, the real distinction is architectural: an execution layer that returns an answer, versus evidence and event data that remain available to the agent after the answer. ## Delegated research tasks and durable source-event data The competitor column summarizes what the named vendor's cited first-party materials describe; the NOSIBLE column summarizes NOSIBLE's cited materials. Primary use NOSIBLE Dated source and event intelligence [[ 4 a]](https://archive.is/YwGWR) Parallel Web Systems Web research through search and task APIs [[ 2 a]](https://archive.is/Rltif) [[ 2 b]](https://web.archive.org/web/20260725105732/https://docs.parallel.ai/task-api/task-quickstart) Search interaction NOSIBLE Agent-oriented source retrieval [[ 4 a]](https://archive.is/YwGWR) Parallel Web Systems Natural-language objective with LLM-optimized excerpts [[ 1 a]](https://archive.is/cN9EP) [[ 1 b]](https://web.archive.org/web/20260725105649/https://docs.parallel.ai/search/search-quickstart) Research execution NOSIBLE Sources and ranked events for downstream workflows [[ 5 a]](https://archive.is/fv0cj) Parallel Web Systems Task API combines inference, search, live crawling, and structured output [[ 2 a]](https://archive.is/Rltif) [[ 2 b]](https://web.archive.org/web/20260725105732/https://docs.parallel.ai/task-api/task-quickstart) Citations NOSIBLE Source-attributed retrieval and event context Parallel Web Systems Task responses include citations [[ 2 a]](https://archive.is/Rltif) [[ 2 b]](https://web.archive.org/web/20260725105732/https://docs.parallel.ai/task-api/task-quickstart) Crawling NOSIBLE Open-web retrieval within NOSIBLE products [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Parallel Web Systems Live crawling is part of Parallel's published Task API workflow [[ 2 a]](https://archive.is/Rltif) [[ 2 b]](https://web.archive.org/web/20260725105732/https://docs.parallel.ai/task-api/task-quickstart) Point-in-time method NOSIBLE Documented buyer evaluation Parallel Web Systems Evaluate historical behavior with a documented buyer test Event representation NOSIBLE Ranked dated WORLD events [[ 5 a]](https://archive.is/fv0cj) Parallel Web Systems Task-specific research output Delivery NOSIBLE API, SDKs, MCP, SEARCH, and WORLD [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) Parallel Web Systems Search API and Task API [[ 1 a]](https://archive.is/cN9EP) [[ 1 b]](https://web.archive.org/web/20260725105649/https://docs.parallel.ai/search/search-quickstart) [[ 2 a]](https://archive.is/Rltif) [[ 2 b]](https://web.archive.org/web/20260725105732/https://docs.parallel.ai/task-api/task-quickstart) Pricing NOSIBLE Confirm current terms with NOSIBLE Parallel Web Systems Confirm current terms with Parallel | Dimension | NOSIBLE | Parallel Web Systems | | --- | --- | --- | | Primary use | Dated source and event intelligence [[ 4 a]](https://archive.is/YwGWR) | Web research through search and task APIs [[ 2 a]](https://archive.is/Rltif) [[ 2 b]](https://web.archive.org/web/20260725105732/https://docs.parallel.ai/task-api/task-quickstart) | | Search interaction | Agent-oriented source retrieval [[ 4 a]](https://archive.is/YwGWR) | Natural-language objective with LLM-optimized excerpts [[ 1 a]](https://archive.is/cN9EP) [[ 1 b]](https://web.archive.org/web/20260725105649/https://docs.parallel.ai/search/search-quickstart) | | Research execution | Sources and ranked events for downstream workflows [[ 5 a]](https://archive.is/fv0cj) | Task API combines inference, search, live crawling, and structured output [[ 2 a]](https://archive.is/Rltif) [[ 2 b]](https://web.archive.org/web/20260725105732/https://docs.parallel.ai/task-api/task-quickstart) | | Citations | Source-attributed retrieval and event context | Task responses include citations [[ 2 a]](https://archive.is/Rltif) [[ 2 b]](https://web.archive.org/web/20260725105732/https://docs.parallel.ai/task-api/task-quickstart) | | Crawling | Open-web retrieval within NOSIBLE products [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Live crawling is part of Parallel's published Task API workflow [[ 2 a]](https://archive.is/Rltif) [[ 2 b]](https://web.archive.org/web/20260725105732/https://docs.parallel.ai/task-api/task-quickstart) | | Point-in-time method | Documented buyer evaluation | Evaluate historical behavior with a documented buyer test | | Event representation | Ranked dated WORLD events [[ 5 a]](https://archive.is/fv0cj) | Task-specific research output | | Delivery | API, SDKs, MCP, SEARCH, and WORLD [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) | Search API and Task API [[ 1 a]](https://archive.is/cN9EP) [[ 1 b]](https://web.archive.org/web/20260725105649/https://docs.parallel.ai/search/search-quickstart) [[ 2 a]](https://archive.is/Rltif) [[ 2 b]](https://web.archive.org/web/20260725105732/https://docs.parallel.ai/task-api/task-quickstart) | | Pricing | Confirm current terms with NOSIBLE | Confirm current terms with Parallel | ## One-call research delivery versus retained evidence In NOSIBLE's view, Parallel may be the more direct choice when a product needs one API call to research an objective and return a structured cited result. NOSIBLE emphasizes dated retrieval, ranked events, and an embedding per event for downstream analysis. [[ 5 a]](https://archive.is/fv0cj) [[ 8 a]](https://archive.is/1CUvQ) In NOSIBLE's view, teams that need both can assess Parallel as an execution layer and NOSIBLE as the evidence layer beneath their own workflow. ## Common Parallel Web Systems comparison questions ### How does NOSIBLE feel it differentiates itself from Parallel Web Systems? NOSIBLE is an AI-native company with two products: SEARCH and WORLD. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 5 a]](https://archive.is/fv0cj) [[ 7 a]](https://archive.is/iDZeo) SEARCH lets agents find dated open-web sources they can cite and inspect directly. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 4 a]](https://archive.is/YwGWR) WORLD is a live open-web event database for models and backtests, with an embedding per event. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 5 a]](https://archive.is/fv0cj) [[ 6 a]](https://archive.is/tnkpG) [[ 8 a]](https://archive.is/1CUvQ) NOSIBLE is committed to open-source software and makes its models publicly available on Hugging Face. [[ 9 a]](https://archive.is/7SSz0) [[ 10 a]](https://archive.is/kHxMG) Related [WORLD v1.2 trial](https://nosible.com/start-trial#data-coverage) [Sentiment model](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) [Forward-looking model](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) ### What does Parallel's Task API do? Parallel says its Task API combines AI inference, web search, live crawling, and structured output with citations. [[ 2 a]](https://archive.is/Rltif) [[ 2 b]](https://web.archive.org/web/20260725105732/https://docs.parallel.ai/task-api/task-quickstart) NOSIBLE publishes dated source retrieval and ranked events rather than a general delegated-task endpoint. [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) In NOSIBLE's view, Parallel may be the more direct choice when an application needs a cited web-research task executed through one request. Related [Agentic Search](https://docs.nosible.com/) [WORLD event database](https://nosible.world/world) ### How does Parallel Search differ from NOSIBLE SEARCH? Parallel documents a Search API that accepts a natural-language objective and returns LLM-optimized excerpts. [[ 1 a]](https://archive.is/cN9EP) [[ 1 b]](https://web.archive.org/web/20260725105649/https://docs.parallel.ai/search/search-quickstart) NOSIBLE SEARCH emphasizes dated open-web sources for agents and research workflows. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 4 a]](https://archive.is/YwGWR) In NOSIBLE's view, buyers should test query control, sources, dates, output structure, and downstream evaluation needs rather than treat the APIs as interchangeable. Related [Agentic Search](https://docs.nosible.com/) [WORLD event database](https://nosible.world/world) ### Can Parallel and NOSIBLE work together? Potentially. Parallel can provide a research-task or search layer, while NOSIBLE can provide dated source retrieval and ranked event history. [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) In NOSIBLE's view, teams should keep provenance fields and test whether citations, dates, and data rights remain suitable before combining results in a production workflow. Related [WORLD event database](https://nosible.world/world) ### How should a team assess citation quality? Use a fixed evaluation set with answers that require multiple sources, changing facts, and explicit dates. Parallel publishes cited task outputs, while NOSIBLE emphasizes source-attributed retrieval. Review citation relevance, source access, date handling, completeness, and the ability to inspect the underlying material before choosing a production dependency. Related [WORLD event database](https://nosible.world/world) ### Does Parallel publish a comparable event database? The cited Parallel documentation describes search objectives and task execution, not a published ranked event database comparable to NOSIBLE WORLD. [[ 5 a]](https://archive.is/fv0cj) This page therefore does not compare aggregate event counts. In NOSIBLE's view, buyers should compare task-level outputs and data definitions before making an equivalence claim. Related [WORLD event database](https://nosible.world/world) Continue comparing ## Related Vendor Comparisons Compare Parallel Web Systems with adjacent options for task execution, agent tooling, data access, and research workflows, then evaluate the output shape your application actually needs for production. [Exa](https://nosible.com/compare/nosible-vs-exa) [Tavily](https://nosible.com/compare/nosible-vs-tavily) [News API](https://nosible.com/compare/nosible-vs-news-api) [Common Crawl](https://nosible.com/compare/nosible-vs-common-crawl) Dive deeper ## Take your next step today Review the delivered field definitions, classification boundaries, and example values before comparing vendor workflows, so your team can assess what each product actually returns to downstream systems. [Data dictionaries](https://nosible.com/data-dictionaries) [Ontology reference](https://nosible.com/ontologies) [API reference](https://docs.nosible.com/) [Start trial](https://nosible.com/start-trial) Sources reviewed July 25, 2026 : [[ 1 ] Parallel Search API quickstart](https://docs.parallel.ai/search/search-quickstart) ( [dated snapshot](https://archive.is/cN9EP); [Wayback copy](https://web.archive.org/web/20260725105649/https://docs.parallel.ai/search/search-quickstart) ) , [[ 2 ] Parallel Task API quickstart](https://docs.parallel.ai/task-api/task-quickstart) ( [dated snapshot](https://archive.is/Rltif); [Wayback copy](https://web.archive.org/web/20260725105732/https://docs.parallel.ai/task-api/task-quickstart) ) , [[ 3 ] NOSIBLE product overview](https://nosible.com/) ( [dated snapshot](https://archive.is/8abI7); [Wayback copy](https://web.archive.org/web/20260713185114/https://nosible.com/) ) , [[ 4 ] NOSIBLE SEARCH](https://docs.nosible.com/) ( [dated snapshot](https://archive.is/YwGWR) ) , [[ 5 ] NOSIBLE WORLD](https://nosible.world/world) ( [dated snapshot](https://archive.is/fv0cj) ) , [[ 6 ] NOSIBLE WORLD v1.2 trial and coverage](https://nosible.com/start-trial) ( [dated snapshot](https://archive.is/tnkpG) ) , [[ 7 ] NOSIBLE AI-native research overview](https://nosible.com/blog) ( [dated snapshot](https://archive.is/iDZeo) ) , [[ 8 ] NOSIBLE embedding-based research](https://nosible.com/blog/an-embedding-based-approach-to-trade-and-economic-policy-uncertainty) ( [dated snapshot](https://archive.is/1CUvQ) ) , [[ 9 ] NOSIBLE Financial Sentiment v1.2 Base](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) ( [dated snapshot](https://archive.is/7SSz0) ) , [[ 10 ] NOSIBLE Forward-Looking v1.2 Base](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) ( [dated snapshot](https://archive.is/kHxMG) ) . Parallel Web Systems is used solely to identify the compared product. NOSIBLE is not affiliated with, sponsored by, or endorsed by Parallel Web Systems. Comparison standards, legal context & corrections This comparison was prepared by NOSIBLE, which has a commercial interest in the products being compared. It is based on the cited public materials as they appeared on July 25, 2026. NOSIBLE has not tested every competitor feature, and the page is not a complete statement of either product. Products change; confirm current requirements, availability, and commercial terms with each vendor. Factual claims are attributed to cited first-party materials; evaluative statements reflect NOSIBLE's opinion. Each source includes a dated archive.is snapshot and, where the Internet Archive captured that URL, a timestamp-specific Wayback copy. Archive availability is controlled by those services. Competitor names and marks are used only to identify the products being compared. No affiliation, sponsorship, or endorsement is implied. The FTC says truthful, non-deceptive comparative advertising may identify competitors ( [policy](https://www.ftc.gov/legal-library/browse/statement-policy-regarding-comparative-advertising); [dated snapshot](https://archive.is/nZXY4); [Wayback copy](https://web.archive.org/web/20260215042003/https://www.ftc.gov/legal-library/browse/statement-policy-regarding-comparative-advertising) ). For one U.S. example of nominative-use analysis, see *New Kids on the Block v. News America Publishing, Inc.*, 971 F.2d 302, 308 (9th Cir. 1992) ( [opinion](https://law.justia.com/cases/federal/appellate-courts/F2/971/302/72076/); [dated snapshot](https://archive.is/dnk8j); [Wayback copy](https://web.archive.org/web/20250211213616/https://law.justia.com/cases/federal/appellate-courts/F2/971/302/72076/) ). If you represent Parallel Web Systems and believe a factual statement is inaccurate, email [stuart@nosible.com](mailto:stuart@nosible.com?subject=Correction%20request%3A%20NOSIBLE%20vs%20Parallel%20Web%20Systems) with the specific claim and a supporting first-party URL. NOSIBLE will review and correct substantiated errors. > Compare NOSIBLE and Parallel Web Systems for web research APIs, task execution, source evidence, citations, event history, and agent workflows. **URL:** https://nosible.com/compare/nosible-vs-parallel --- --- title: "NOSIBLE vs Tavily" description: "Compare NOSIBLE and Tavily for agent search APIs, extraction, crawling, research, source evidence, event history, and delivery options." url: "https://nosible.com/compare/nosible-vs-tavily" --- Comparison / Reviewed July 25, 2026 # NOSIBLE vs Tavily Tavily documents a modular toolbox for agent search, extraction, crawling, mapping, and research. [[ 1 a]](https://archive.is/8Utdj) NOSIBLE gives agents a narrower but deeper research substrate: dated sources and ranked events. [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) NOSIBLE-AUTHORED COMPARISON · NOSIBLE HAS A COMMERCIAL INTEREST IN THIS COMPARISON · FACTS ATTRIBUTED TO FIRST-PARTY VENDOR MATERIALS · EVALUATIVE STATEMENTS ARE NOSIBLE'S OPINION · REVIEWED JULY 25, 2026 · DUAL PUBLIC ARCHIVES WHERE SUPPORTED · [STANDARDS & CORRECTIONS](https://nosible.com/compare/nosible-vs-tavily#comparison-standards). IF YOU REPRESENT TAVILY AND BELIEVE A FACTUAL STATEMENT IS INACCURATE, EMAIL [STUART@NOSIBLE.COM](mailto:stuart@nosible.com?subject=Correction%20request%3A%20NOSIBLE%20vs%20Tavily) WITH THE SPECIFIC CLAIM AND A SUPPORTING FIRST-PARTY URL. NOSIBLE WILL REVIEW AND CORRECT SUBSTANTIATED ERRORS. - Tavily documents APIs for search, extract, crawl, map, and research workflows built for AI applications. [[ 1 a]](https://archive.is/8Utdj) - Tavily's Search endpoint accepts a query and configurable search depth. [[ 1 a]](https://archive.is/8Utdj) [[ 2 a]](https://archive.is/mLNqG) - NOSIBLE SEARCH focuses on dated open-web source evidence, while WORLD supplies ranked event history with an embedding per event. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 5 a]](https://archive.is/fv0cj) [[ 8 a]](https://archive.is/1CUvQ) - In NOSIBLE's view, Tavily may be the more direct fit for applications that need its modular web-search and crawling endpoints. - In NOSIBLE's view, NOSIBLE is the stronger fit when dated event context and replayable evidence are core requirements. ## A web-toolbox decision, not a feature-count contest Tavily documents separate APIs for search, extraction, crawling, mapping, and research in AI applications. [[ 1 a]](https://archive.is/8Utdj) NOSIBLE deliberately offers a more opinionated research substrate: dated open-web sources in SEARCH and ranked events in WORLD. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 5 a]](https://archive.is/fv0cj) In NOSIBLE's view, the useful question is whether an agent needs to operate on the web, or needs a source-and-event record it can inspect and analyse over time. ## Agent web tooling and an inspectable research substrate The competitor column summarizes what the named vendor's cited first-party materials describe; the NOSIBLE column summarizes NOSIBLE's cited materials. Primary use NOSIBLE Dated source and event intelligence [[ 4 a]](https://archive.is/YwGWR) Tavily Modular web-search and research APIs for AI applications [[ 1 a]](https://archive.is/8Utdj) Search NOSIBLE Agent-oriented dated source retrieval [[ 4 a]](https://archive.is/YwGWR) Tavily Search endpoint with query and search-depth controls [[ 1 a]](https://archive.is/8Utdj) [[ 2 a]](https://archive.is/mLNqG) Extraction NOSIBLE Source retrieval within SEARCH workflows [[ 4 a]](https://archive.is/YwGWR) Tavily Dedicated Extract endpoint [[ 1 a]](https://archive.is/8Utdj) Crawling and mapping NOSIBLE Open-web retrieval within NOSIBLE products [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Tavily Dedicated Crawl and Map endpoints [[ 1 a]](https://archive.is/8Utdj) Research NOSIBLE Sources and WORLD events for downstream research [[ 5 a]](https://archive.is/fv0cj) Tavily Dedicated Research API [[ 1 a]](https://archive.is/8Utdj) Point-in-time method NOSIBLE Documented buyer evaluation Tavily Evaluate historical behavior using a documented buyer test Event representation NOSIBLE Ranked dated WORLD events [[ 5 a]](https://archive.is/fv0cj) Tavily Search and research results rather than a published event database [[ 1 a]](https://archive.is/8Utdj) Delivery NOSIBLE API, SDKs, MCP, SEARCH, and WORLD [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) Tavily Search, extract, crawl, map, and research endpoints [[ 1 a]](https://archive.is/8Utdj) [[ 2 a]](https://archive.is/mLNqG) Pricing NOSIBLE Confirm current terms with NOSIBLE Tavily Confirm current terms with Tavily | Dimension | NOSIBLE | Tavily | | --- | --- | --- | | Primary use | Dated source and event intelligence [[ 4 a]](https://archive.is/YwGWR) | Modular web-search and research APIs for AI applications [[ 1 a]](https://archive.is/8Utdj) | | Search | Agent-oriented dated source retrieval [[ 4 a]](https://archive.is/YwGWR) | Search endpoint with query and search-depth controls [[ 1 a]](https://archive.is/8Utdj) [[ 2 a]](https://archive.is/mLNqG) | | Extraction | Source retrieval within SEARCH workflows [[ 4 a]](https://archive.is/YwGWR) | Dedicated Extract endpoint [[ 1 a]](https://archive.is/8Utdj) | | Crawling and mapping | Open-web retrieval within NOSIBLE products [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Dedicated Crawl and Map endpoints [[ 1 a]](https://archive.is/8Utdj) | | Research | Sources and WORLD events for downstream research [[ 5 a]](https://archive.is/fv0cj) | Dedicated Research API [[ 1 a]](https://archive.is/8Utdj) | | Point-in-time method | Documented buyer evaluation | Evaluate historical behavior using a documented buyer test | | Event representation | Ranked dated WORLD events [[ 5 a]](https://archive.is/fv0cj) | Search and research results rather than a published event database [[ 1 a]](https://archive.is/8Utdj) | | Delivery | API, SDKs, MCP, SEARCH, and WORLD [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) | Search, extract, crawl, map, and research endpoints [[ 1 a]](https://archive.is/8Utdj) [[ 2 a]](https://archive.is/mLNqG) | | Pricing | Confirm current terms with NOSIBLE | Confirm current terms with Tavily | ## Web operations versus evidence-led research In NOSIBLE's view, Tavily may be the more direct choice when an application needs its discrete search, extract, crawl, map, or research endpoints. NOSIBLE is designed for agents that need dated source evidence and ranked events, with an embedding per event for downstream use. [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) [[ 8 a]](https://archive.is/1CUvQ) In NOSIBLE's view, both deserve a task-level trial when web operations must feed a historical research workflow. ## Common Tavily comparison questions ### How does NOSIBLE feel it differentiates itself from Tavily? NOSIBLE is an AI-native company with two products: SEARCH and WORLD. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 5 a]](https://archive.is/fv0cj) [[ 7 a]](https://archive.is/iDZeo) SEARCH lets agents find dated open-web sources they can cite and inspect directly. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 4 a]](https://archive.is/YwGWR) WORLD is a live open-web event database for models and backtests, with an embedding per event. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 5 a]](https://archive.is/fv0cj) [[ 6 a]](https://archive.is/tnkpG) [[ 8 a]](https://archive.is/1CUvQ) NOSIBLE is committed to open-source software and makes its models publicly available on Hugging Face. [[ 9 a]](https://archive.is/7SSz0) [[ 10 a]](https://archive.is/kHxMG) Related [WORLD v1.2 trial](https://nosible.com/start-trial#data-coverage) [Sentiment model](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) [Forward-looking model](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) ### What APIs does Tavily publish? Tavily's documentation lists APIs for search, extraction, crawling, mapping, and research. [[ 1 a]](https://archive.is/8Utdj) Its Search endpoint accepts a query and configurable search depth. [[ 1 a]](https://archive.is/8Utdj) [[ 2 a]](https://archive.is/mLNqG) NOSIBLE publishes SEARCH and WORLD products for dated source and event workflows. [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) In NOSIBLE's view, Tavily may be the more direct fit when an application specifically needs its modular endpoint set. Related [Agentic Search](https://docs.nosible.com/) [WORLD event database](https://nosible.world/world) ### How does Tavily Search differ from NOSIBLE SEARCH? Tavily documents a configurable search endpoint for AI applications. [[ 1 a]](https://archive.is/8Utdj) [[ 2 a]](https://archive.is/mLNqG) NOSIBLE SEARCH emphasizes dated open-web sources for agents and research. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 4 a]](https://archive.is/YwGWR) In NOSIBLE's view, buyers should compare the exact queries, source types, extraction needs, date fields, delivery method, and evaluation criteria that their application requires before selecting either product. Related [Agentic Search](https://docs.nosible.com/) [WORLD event database](https://nosible.world/world) ### Can Tavily and NOSIBLE be complementary? Potentially. Tavily can provide search, extract, crawl, map, and research endpoints, while NOSIBLE can supply dated source evidence and ranked events. [[ 1 a]](https://archive.is/8Utdj) [[ 2 a]](https://archive.is/mLNqG) [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) In NOSIBLE's view, teams should test source provenance, duplication, timing, and rights before combining results from separate web-data services in a production system. Related [Agentic Search](https://docs.nosible.com/) [WORLD event database](https://nosible.world/world) ### Does Tavily publish a comparable event database? The cited Tavily materials describe web-search, extraction, crawling, mapping, and research endpoints, not a published ranked event database comparable to NOSIBLE WORLD. [[ 1 a]](https://archive.is/8Utdj) [[ 2 a]](https://archive.is/mLNqG) [[ 5 a]](https://archive.is/fv0cj) This page therefore does not compare aggregate event counts. In NOSIBLE's view, output definitions should be tested at the workflow level rather than assumed to be equivalent. Related [Agentic Search](https://docs.nosible.com/) [WORLD event database](https://nosible.world/world) ### How should a team evaluate web-crawling tools? Create a documented sample of allowed domains, public pages, difficult layouts, and historical questions. Measure access success, extraction quality, rendered content, dates, citations, latency, and governance controls. [[ 1 a]](https://archive.is/8Utdj) The right choice depends on the sources, tasks, and evidence records needed by the deployment, rather than a generic feature checklist. Related [WORLD event database](https://nosible.world/world) [HTML to JSON](https://docs.nosible.com/endpoints/search-scrape) Continue comparing ## Related Vendor Comparisons Compare Tavily with adjacent options for task execution, agent tooling, data access, and research workflows, then evaluate the output shape your application actually needs for production. [Exa](https://nosible.com/compare/nosible-vs-exa) [Parallel](https://nosible.com/compare/nosible-vs-parallel) [Common Crawl](https://nosible.com/compare/nosible-vs-common-crawl) [News API](https://nosible.com/compare/nosible-vs-news-api) Dive deeper ## Take your next step today Review the delivered field definitions, classification boundaries, and example values before comparing vendor workflows, so your team can assess what each product actually returns to downstream systems. [Data dictionaries](https://nosible.com/data-dictionaries) [Ontology reference](https://nosible.com/ontologies) [API reference](https://docs.nosible.com/) [Start trial](https://nosible.com/start-trial) Sources reviewed July 25, 2026 : [[ 1 ] Tavily API introduction](https://docs.tavily.com/documentation/api-reference/introduction) ( [dated snapshot](https://archive.is/8Utdj) ) , [[ 2 ] Tavily Search API](https://docs.tavily.com/documentation/api-reference/endpoint/search) ( [dated snapshot](https://archive.is/mLNqG) ) , [[ 3 ] NOSIBLE product overview](https://nosible.com/) ( [dated snapshot](https://archive.is/8abI7); [Wayback copy](https://web.archive.org/web/20260713185114/https://nosible.com/) ) , [[ 4 ] NOSIBLE SEARCH](https://docs.nosible.com/) ( [dated snapshot](https://archive.is/YwGWR) ) , [[ 5 ] NOSIBLE WORLD](https://nosible.world/world) ( [dated snapshot](https://archive.is/fv0cj) ) , [[ 6 ] NOSIBLE WORLD v1.2 trial and coverage](https://nosible.com/start-trial) ( [dated snapshot](https://archive.is/tnkpG) ) , [[ 7 ] NOSIBLE AI-native research overview](https://nosible.com/blog) ( [dated snapshot](https://archive.is/iDZeo) ) , [[ 8 ] NOSIBLE embedding-based research](https://nosible.com/blog/an-embedding-based-approach-to-trade-and-economic-policy-uncertainty) ( [dated snapshot](https://archive.is/1CUvQ) ) , [[ 9 ] NOSIBLE Financial Sentiment v1.2 Base](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) ( [dated snapshot](https://archive.is/7SSz0) ) , [[ 10 ] NOSIBLE Forward-Looking v1.2 Base](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) ( [dated snapshot](https://archive.is/kHxMG) ) . Tavily is used solely to identify the compared product. NOSIBLE is not affiliated with, sponsored by, or endorsed by Tavily. Comparison standards, legal context & corrections This comparison was prepared by NOSIBLE, which has a commercial interest in the products being compared. It is based on the cited public materials as they appeared on July 25, 2026. NOSIBLE has not tested every competitor feature, and the page is not a complete statement of either product. Products change; confirm current requirements, availability, and commercial terms with each vendor. Factual claims are attributed to cited first-party materials; evaluative statements reflect NOSIBLE's opinion. Each source includes a dated archive.is snapshot and, where the Internet Archive captured that URL, a timestamp-specific Wayback copy. Archive availability is controlled by those services. Competitor names and marks are used only to identify the products being compared. No affiliation, sponsorship, or endorsement is implied. The FTC says truthful, non-deceptive comparative advertising may identify competitors ( [policy](https://www.ftc.gov/legal-library/browse/statement-policy-regarding-comparative-advertising); [dated snapshot](https://archive.is/nZXY4); [Wayback copy](https://web.archive.org/web/20260215042003/https://www.ftc.gov/legal-library/browse/statement-policy-regarding-comparative-advertising) ). For one U.S. example of nominative-use analysis, see *New Kids on the Block v. News America Publishing, Inc.*, 971 F.2d 302, 308 (9th Cir. 1992) ( [opinion](https://law.justia.com/cases/federal/appellate-courts/F2/971/302/72076/); [dated snapshot](https://archive.is/dnk8j); [Wayback copy](https://web.archive.org/web/20250211213616/https://law.justia.com/cases/federal/appellate-courts/F2/971/302/72076/) ). If you represent Tavily and believe a factual statement is inaccurate, email [stuart@nosible.com](mailto:stuart@nosible.com?subject=Correction%20request%3A%20NOSIBLE%20vs%20Tavily) with the specific claim and a supporting first-party URL. NOSIBLE will review and correct substantiated errors. > Compare NOSIBLE and Tavily for agent search APIs, extraction, crawling, research, source evidence, event history, and delivery options. **URL:** https://nosible.com/compare/nosible-vs-tavily --- --- title: "NOSIBLE vs News API" description: "Compare NOSIBLE and News API for news search, headlines, source coverage, response content, dated evidence, event history, and research workflows." url: "https://nosible.com/compare/nosible-vs-news-api" --- Comparison / Reviewed July 25, 2026 # NOSIBLE vs News API News API documents straightforward endpoints for article discovery and breaking headlines. [[ 1 a]](https://archive.is/FQqHw) [[ 2 a]](https://archive.is/J04Ki) NOSIBLE carries the research workflow further, pairing dated source evidence with ranked events for agents and historical analysis. [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) NOSIBLE-AUTHORED COMPARISON · NOSIBLE HAS A COMMERCIAL INTEREST IN THIS COMPARISON · FACTS ATTRIBUTED TO FIRST-PARTY VENDOR MATERIALS · EVALUATIVE STATEMENTS ARE NOSIBLE'S OPINION · REVIEWED JULY 25, 2026 · DUAL PUBLIC ARCHIVES WHERE SUPPORTED · [STANDARDS & CORRECTIONS](https://nosible.com/compare/nosible-vs-news-api#comparison-standards). IF YOU REPRESENT NEWS API AND BELIEVE A FACTUAL STATEMENT IS INACCURATE, EMAIL [STUART@NOSIBLE.COM](mailto:stuart@nosible.com?subject=Correction%20request%3A%20NOSIBLE%20vs%20News%20API) WITH THE SPECIFIC CLAIM AND A SUPPORTING FIRST-PARTY URL. NOSIBLE WILL REVIEW AND CORRECT SUBSTANTIATED ERRORS. - News API documents an Everything endpoint for article discovery and a Top Headlines endpoint for breaking headlines. [[ 1 a]](https://archive.is/FQqHw) [[ 2 a]](https://archive.is/J04Ki) - News API says Everything covers more than 150,000 sources over the past five years. [[ 1 a]](https://archive.is/FQqHw) [[ 2 a]](https://archive.is/J04Ki) - News API documents that article content is truncated to 200 characters in its Everything responses. [[ 1 a]](https://archive.is/FQqHw) [[ 2 a]](https://archive.is/J04Ki) - NOSIBLE SEARCH and WORLD turn dated open-web evidence into a source-and-event research layer rather than a two-endpoint news feed. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 5 a]](https://archive.is/fv0cj) - In NOSIBLE's view, NOSIBLE is the stronger fit when source inspection and event-oriented historical analysis are requirements. ## From a news result to an evidence-backed event News API documents endpoints for article discovery and breaking headlines, including a published source-coverage claim for Everything. [[ 1 a]](https://archive.is/FQqHw) [[ 2 a]](https://archive.is/J04Ki) NOSIBLE uses a different unit of work: dated source evidence in SEARCH and ranked events in WORLD. [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) In NOSIBLE's view, that makes NOSIBLE more relevant when a product must do more than surface a headline—it must inspect the source, preserve timing, and analyse an event over time. ## Headline delivery and source-to-event research The competitor column summarizes what the named vendor's cited first-party materials describe; the NOSIBLE column summarizes NOSIBLE's cited materials. Primary use NOSIBLE Dated source and event intelligence [[ 4 a]](https://archive.is/YwGWR) News API News discovery and breaking-headline API access [[ 2 a]](https://archive.is/J04Ki) Published endpoints NOSIBLE SEARCH and WORLD products [[ 5 a]](https://archive.is/fv0cj) News API Everything and Top Headlines endpoints [[ 1 a]](https://archive.is/FQqHw) [[ 2 a]](https://archive.is/J04Ki) Article discovery NOSIBLE Open-web sources with dated evidence [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) News API Everything endpoint for discovery and analysis [[ 1 a]](https://archive.is/FQqHw) [[ 2 a]](https://archive.is/J04Ki) Headlines NOSIBLE Source and event retrieval News API Top Headlines endpoint for breaking headlines [[ 1 a]](https://archive.is/FQqHw) Published source scope NOSIBLE Test source scope for the task News API News API says 150k+ sources over the past five years [[ 1 a]](https://archive.is/FQqHw) Content in response NOSIBLE Inspectable source retrieval [[ 4 a]](https://archive.is/YwGWR) News API Everything article content documented as truncated to 200 characters [[ 1 a]](https://archive.is/FQqHw) [[ 2 a]](https://archive.is/J04Ki) Point-in-time method NOSIBLE Documented buyer evaluation News API Evaluate retention and historical behavior with a buyer test Event representation NOSIBLE Ranked dated WORLD events [[ 5 a]](https://archive.is/fv0cj) News API News article and headline results Pricing NOSIBLE Confirm current terms with NOSIBLE News API Confirm current terms with News API | Dimension | NOSIBLE | News API | | --- | --- | --- | | Primary use | Dated source and event intelligence [[ 4 a]](https://archive.is/YwGWR) | News discovery and breaking-headline API access [[ 2 a]](https://archive.is/J04Ki) | | Published endpoints | SEARCH and WORLD products [[ 5 a]](https://archive.is/fv0cj) | Everything and Top Headlines endpoints [[ 1 a]](https://archive.is/FQqHw) [[ 2 a]](https://archive.is/J04Ki) | | Article discovery | Open-web sources with dated evidence [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Everything endpoint for discovery and analysis [[ 1 a]](https://archive.is/FQqHw) [[ 2 a]](https://archive.is/J04Ki) | | Headlines | Source and event retrieval | Top Headlines endpoint for breaking headlines [[ 1 a]](https://archive.is/FQqHw) | | Published source scope | Test source scope for the task | News API says 150k+ sources over the past five years [[ 1 a]](https://archive.is/FQqHw) | | Content in response | Inspectable source retrieval [[ 4 a]](https://archive.is/YwGWR) | Everything article content documented as truncated to 200 characters [[ 1 a]](https://archive.is/FQqHw) [[ 2 a]](https://archive.is/J04Ki) | | Point-in-time method | Documented buyer evaluation | Evaluate retention and historical behavior with a buyer test | | Event representation | Ranked dated WORLD events [[ 5 a]](https://archive.is/fv0cj) | News article and headline results | | Pricing | Confirm current terms with NOSIBLE | Confirm current terms with News API | ## A clean news feed versus a research layer In NOSIBLE's view, News API may be the more direct choice when an application architecture calls for its Everything or Top Headlines endpoint pattern. NOSIBLE is designed for dated sources and ranked events, with an embedding per event for downstream models and backtests. [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) [[ 8 a]](https://archive.is/1CUvQ) In NOSIBLE's view, buyers should test both against the point at which a news result must become research evidence. ## Common News API comparison questions ### How does NOSIBLE feel it differentiates itself from News API? NOSIBLE is an AI-native company with two products: SEARCH and WORLD. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 5 a]](https://archive.is/fv0cj) [[ 7 a]](https://archive.is/iDZeo) SEARCH lets agents find dated open-web sources they can cite and inspect directly. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 4 a]](https://archive.is/YwGWR) WORLD is a live open-web event database for models and backtests, with an embedding per event. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 5 a]](https://archive.is/fv0cj) [[ 6 a]](https://archive.is/tnkpG) [[ 8 a]](https://archive.is/1CUvQ) NOSIBLE is committed to open-source software and makes its models publicly available on Hugging Face. [[ 9 a]](https://archive.is/7SSz0) [[ 10 a]](https://archive.is/kHxMG) Related [WORLD v1.2 trial](https://nosible.com/start-trial#data-coverage) [Sentiment model](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) [Forward-looking model](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) ### What endpoints does News API publish? News API documents an Everything endpoint for article discovery and analysis and a Top Headlines endpoint for breaking headlines. [[ 1 a]](https://archive.is/FQqHw) [[ 2 a]](https://archive.is/J04Ki) NOSIBLE publishes SEARCH and WORLD products for dated source and event workflows. [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) In NOSIBLE's view, News API may be the more direct fit when those two news-feed endpoints match the application architecture. Related [Agentic Search](https://docs.nosible.com/) [WORLD event database](https://nosible.world/world) ### What does News API say about coverage and response content? News API says its Everything endpoint covers more than 150,000 sources over the past five years, and its documentation says article content is truncated to 200 characters. [[ 1 a]](https://archive.is/FQqHw) [[ 2 a]](https://archive.is/J04Ki) Those vendor-reported details should be tested against the required publishers, retention, extraction needs, and rights before a buyer selects a production source. Related [Agentic Search](https://docs.nosible.com/) [WORLD event database](https://nosible.world/world) ### How does NOSIBLE differ for historical research? NOSIBLE emphasizes dated open-web sources and ranked events for research agents and historical analysis. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) News API documents article discovery and headline retrieval. [[ 2 a]](https://archive.is/J04Ki) In NOSIBLE's view, teams should run fixed historical queries and record source URLs, available content, publication times, retention, revisions, and the fields needed for the downstream model or review process. Related [Agentic Search](https://docs.nosible.com/) [WORLD event database](https://nosible.world/world) ### Does News API publish a comparable event database? The cited News API documentation describes article-discovery and headline endpoints, not a published ranked event database comparable to NOSIBLE WORLD. [[ 2 a]](https://archive.is/J04Ki) [[ 5 a]](https://archive.is/fv0cj) This page therefore does not compare aggregate event counts. In NOSIBLE's view, article results and event records should be treated as distinct data products until task-level validation shows otherwise. Related [Agentic Search](https://docs.nosible.com/) [WORLD event database](https://nosible.world/world) ### Can News API and NOSIBLE be used together? Potentially. News API can provide article discovery or headline retrieval, while NOSIBLE can provide dated open-web sources and ranked event context. [[ 2 a]](https://archive.is/J04Ki) [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 5 a]](https://archive.is/fv0cj) In NOSIBLE's view, teams should document licensing, retention, duplicates, timestamps, and source provenance before combining outputs in a research, monitoring, or model-building workflow. Related [Agentic Search](https://docs.nosible.com/) [WORLD event database](https://nosible.world/world) Continue comparing ## Related Vendor Comparisons Compare News API with adjacent options for task execution, agent tooling, data access, and research workflows, then evaluate the output shape your application actually needs for production. [Tavily](https://nosible.com/compare/nosible-vs-tavily) [Exa](https://nosible.com/compare/nosible-vs-exa) [Parallel](https://nosible.com/compare/nosible-vs-parallel) [GDELT](https://nosible.com/compare/nosible-vs-gdelt) Dive deeper ## Take your next step today Review the delivered field definitions, classification boundaries, and example values before comparing vendor workflows, so your team can assess what each product actually returns to downstream systems. [Data dictionaries](https://nosible.com/data-dictionaries) [Ontology reference](https://nosible.com/ontologies) [API reference](https://docs.nosible.com/) [Start trial](https://nosible.com/start-trial) Sources reviewed July 25, 2026 : [[ 1 ] News API endpoint documentation](https://newsapi.org/docs/endpoints) ( [dated snapshot](https://archive.is/FQqHw) ) , [[ 2 ] News API Everything endpoint](https://newsapi.org/docs/endpoints/everything) ( [dated snapshot](https://archive.is/J04Ki) ) , [[ 3 ] NOSIBLE product overview](https://nosible.com/) ( [dated snapshot](https://archive.is/8abI7); [Wayback copy](https://web.archive.org/web/20260713185114/https://nosible.com/) ) , [[ 4 ] NOSIBLE SEARCH](https://docs.nosible.com/) ( [dated snapshot](https://archive.is/YwGWR) ) , [[ 5 ] NOSIBLE WORLD](https://nosible.world/world) ( [dated snapshot](https://archive.is/fv0cj) ) , [[ 6 ] NOSIBLE WORLD v1.2 trial and coverage](https://nosible.com/start-trial) ( [dated snapshot](https://archive.is/tnkpG) ) , [[ 7 ] NOSIBLE AI-native research overview](https://nosible.com/blog) ( [dated snapshot](https://archive.is/iDZeo) ) , [[ 8 ] NOSIBLE embedding-based research](https://nosible.com/blog/an-embedding-based-approach-to-trade-and-economic-policy-uncertainty) ( [dated snapshot](https://archive.is/1CUvQ) ) , [[ 9 ] NOSIBLE Financial Sentiment v1.2 Base](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) ( [dated snapshot](https://archive.is/7SSz0) ) , [[ 10 ] NOSIBLE Forward-Looking v1.2 Base](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) ( [dated snapshot](https://archive.is/kHxMG) ) . News API is used solely to identify the compared product. NOSIBLE is not affiliated with, sponsored by, or endorsed by News API. Comparison standards, legal context & corrections This comparison was prepared by NOSIBLE, which has a commercial interest in the products being compared. It is based on the cited public materials as they appeared on July 25, 2026. NOSIBLE has not tested every competitor feature, and the page is not a complete statement of either product. Products change; confirm current requirements, availability, and commercial terms with each vendor. Factual claims are attributed to cited first-party materials; evaluative statements reflect NOSIBLE's opinion. Each source includes a dated archive.is snapshot and, where the Internet Archive captured that URL, a timestamp-specific Wayback copy. Archive availability is controlled by those services. Competitor names and marks are used only to identify the products being compared. No affiliation, sponsorship, or endorsement is implied. The FTC says truthful, non-deceptive comparative advertising may identify competitors ( [policy](https://www.ftc.gov/legal-library/browse/statement-policy-regarding-comparative-advertising); [dated snapshot](https://archive.is/nZXY4); [Wayback copy](https://web.archive.org/web/20260215042003/https://www.ftc.gov/legal-library/browse/statement-policy-regarding-comparative-advertising) ). For one U.S. example of nominative-use analysis, see *New Kids on the Block v. News America Publishing, Inc.*, 971 F.2d 302, 308 (9th Cir. 1992) ( [opinion](https://law.justia.com/cases/federal/appellate-courts/F2/971/302/72076/); [dated snapshot](https://archive.is/dnk8j); [Wayback copy](https://web.archive.org/web/20250211213616/https://law.justia.com/cases/federal/appellate-courts/F2/971/302/72076/) ). If you represent News API and believe a factual statement is inaccurate, email [stuart@nosible.com](mailto:stuart@nosible.com?subject=Correction%20request%3A%20NOSIBLE%20vs%20News%20API) with the specific claim and a supporting first-party URL. NOSIBLE will review and correct substantiated errors. > Compare NOSIBLE and News API for news search, headlines, source coverage, response content, dated evidence, event history, and research workflows. **URL:** https://nosible.com/compare/nosible-vs-news-api --- --- title: "NOSIBLE WORLD vs GDELT" description: "How NOSIBLE WORLD, a web-scale search and market-event intelligence engine for AI agents, compares with GDELT, a free and open global-news research platform. Coverage, history, delivery, and use cases." url: "https://nosible.com/compare/nosible-vs-gdelt" --- Comparison / Reviewed July 14, 2026 # NOSIBLE WORLD vs GDELT GDELT is a free and open platform for studying global society through news, events, people, locations, organizations, themes, and emotions. [[ 2 a]](https://archive.is/IRXzN) [[ 2 b]](https://web.archive.org/web/20260622051800/https://www.gdeltproject.org/data.html) NOSIBLE WORLD is a commercial source-search and market-event layer built for investment, risk, and research agents. [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) NOSIBLE-AUTHORED COMPARISON · NOSIBLE HAS A COMMERCIAL INTEREST IN THIS COMPARISON · FACTS ATTRIBUTED TO FIRST-PARTY VENDOR MATERIALS · EVALUATIVE STATEMENTS ARE NOSIBLE'S OPINION · REVIEWED JULY 14, 2026 · DUAL PUBLIC ARCHIVES WHERE SUPPORTED · [STANDARDS & CORRECTIONS](https://nosible.com/compare/nosible-vs-gdelt#comparison-standards). IF YOU REPRESENT GDELT AND BELIEVE A FACTUAL STATEMENT IS INACCURATE, EMAIL [STUART@NOSIBLE.COM](mailto:stuart@nosible.com?subject=Correction%20request%3A%20NOSIBLE%20WORLD%20vs%20GDELT) WITH THE SPECIFIC CLAIM AND A SUPPORTING FIRST-PARTY URL. NOSIBLE WILL REVIEW AND CORRECT SUBSTANTIATED ERRORS. - GDELT says it monitors print, broadcast, and web news in more than 100 languages and translates 65 languages in real time. [[ 2 a]](https://archive.is/IRXzN) [[ 2 b]](https://web.archive.org/web/20260622051800/https://www.gdeltproject.org/data.html) - GDELT 1.0 includes event data back to January 1, 1979, while GDELT 2.0 datasets update every 15 minutes. [[ 2 a]](https://archive.is/IRXzN) [[ 2 b]](https://web.archive.org/web/20260622051800/https://www.gdeltproject.org/data.html) - GDELT publishes more than 300 event categories and more than 2,200 emotions and themes through its Global Content Analysis Measures. [[ 2 a]](https://archive.is/IRXzN) [[ 2 b]](https://web.archive.org/web/20260622051800/https://www.gdeltproject.org/data.html) - NOSIBLE WORLD emphasizes dated source retrieval, ranked events, ticker context, and managed research-agent workflows. [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) ## Open global research and managed agent evidence GDELT offers open infrastructure for examining the global news ecosystem across decades, languages, events, and themes. [[ 2 a]](https://archive.is/IRXzN) [[ 2 b]](https://web.archive.org/web/20260622051800/https://www.gdeltproject.org/data.html) NOSIBLE WORLD is designed for teams that want dated source retrieval, ranked market events, ticker context, and an API layer for research agents. [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) In NOSIBLE's view, NOSIBLE is the stronger fit when a managed securities-oriented evidence workflow is the defining requirement. ## Feature comparison The competitor column summarizes what the named vendor's cited first-party materials describe; the NOSIBLE column summarizes NOSIBLE's cited materials. Category NOSIBLE Web-scale search and market-event intelligence for AI agents [[ 4 a]](https://archive.is/glqrw) GDELT Free, open global news and event research database [[ 1 a]](https://archive.is/vfGKI) [[ 1 b]](https://web.archive.org/web/20260707152459/https://www.gdeltproject.org/) Content scope NOSIBLE Long-form news, corporate, government; excludes social, paywalled, marketplaces, adult [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) GDELT Print, broadcast, and web news, with event, knowledge-graph, and imagery datasets [[ 1 a]](https://archive.is/vfGKI) [[ 1 b]](https://web.archive.org/web/20260707152459/https://www.gdeltproject.org/) Languages NOSIBLE 95 [[ 4 a]](https://archive.is/glqrw) GDELT 65 machine-translated (100+ monitored) [[ 1 a]](https://archive.is/vfGKI) [[ 1 b]](https://web.archive.org/web/20260707152459/https://www.gdeltproject.org/) History and point-in-time NOSIBLE About 30 years; five-way date-verification method [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) GDELT GDELT 1.0 data from January 1, 1979; GDELT 2.0 updates every 15 minutes [[ 2 a]](https://archive.is/IRXzN) [[ 2 b]](https://web.archive.org/web/20260622051800/https://www.gdeltproject.org/data.html) Event database NOSIBLE Standalone 100M+ dated, ranked events (WORLD) [[ 5 a]](https://archive.is/tGYc9) GDELT More than 300 event categories plus a Global Knowledge Graph [[ 2 a]](https://archive.is/IRXzN) [[ 2 b]](https://web.archive.org/web/20260622051800/https://www.gdeltproject.org/data.html) Sentiment NOSIBLE Available in enrichment [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) GDELT Global Content Analysis Measures covering more than 2,200 emotions and themes [[ 2 a]](https://archive.is/IRXzN) [[ 2 b]](https://web.archive.org/web/20260622051800/https://www.gdeltproject.org/data.html) Structured context NOSIBLE Source-attributed entities and event context [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) GDELT People, locations, organizations, themes, emotions, events, and counts [[ 2 a]](https://archive.is/IRXzN) [[ 2 b]](https://web.archive.org/web/20260622051800/https://www.gdeltproject.org/data.html) Entity and ticker mapping NOSIBLE Source-attributed entities [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) GDELT Global Knowledge Graph entities and themes [[ 2 a]](https://archive.is/IRXzN) [[ 2 b]](https://web.archive.org/web/20260622051800/https://www.gdeltproject.org/data.html) Delivery and API NOSIBLE Six-endpoint agent API, SDKs, MCP [[ 4 a]](https://archive.is/glqrw) GDELT Open datasets, downloads, and analysis tools [[ 1 a]](https://archive.is/vfGKI) [[ 1 b]](https://web.archive.org/web/20260707152459/https://www.gdeltproject.org/) Research orientation NOSIBLE Built for AI-agent and model workflows [[ 4 a]](https://archive.is/glqrw) GDELT Open research into human society and the global news ecosystem [[ 1 a]](https://archive.is/vfGKI) [[ 1 b]](https://web.archive.org/web/20260707152459/https://www.gdeltproject.org/) Typical use NOSIBLE Investment, risk, and research-agent evidence [[ 4 a]](https://archive.is/glqrw) GDELT Global-scale media, event, network, and societal analysis [[ 1 a]](https://archive.is/vfGKI) [[ 1 b]](https://web.archive.org/web/20260707152459/https://www.gdeltproject.org/) Pricing NOSIBLE See nosible.com [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) GDELT Free and open [[ 1 a]](https://archive.is/vfGKI) [[ 1 b]](https://web.archive.org/web/20260707152459/https://www.gdeltproject.org/) | Dimension | NOSIBLE | GDELT | | --- | --- | --- | | Category | Web-scale search and market-event intelligence for AI agents [[ 4 a]](https://archive.is/glqrw) | Free, open global news and event research database [[ 1 a]](https://archive.is/vfGKI) [[ 1 b]](https://web.archive.org/web/20260707152459/https://www.gdeltproject.org/) | | Content scope | Long-form news, corporate, government; excludes social, paywalled, marketplaces, adult [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Print, broadcast, and web news, with event, knowledge-graph, and imagery datasets [[ 1 a]](https://archive.is/vfGKI) [[ 1 b]](https://web.archive.org/web/20260707152459/https://www.gdeltproject.org/) | | Languages | 95 [[ 4 a]](https://archive.is/glqrw) | 65 machine-translated (100+ monitored) [[ 1 a]](https://archive.is/vfGKI) [[ 1 b]](https://web.archive.org/web/20260707152459/https://www.gdeltproject.org/) | | History and point-in-time | About 30 years; five-way date-verification method [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | GDELT 1.0 data from January 1, 1979; GDELT 2.0 updates every 15 minutes [[ 2 a]](https://archive.is/IRXzN) [[ 2 b]](https://web.archive.org/web/20260622051800/https://www.gdeltproject.org/data.html) | | Event database | Standalone 100M+ dated, ranked events (WORLD) [[ 5 a]](https://archive.is/tGYc9) | More than 300 event categories plus a Global Knowledge Graph [[ 2 a]](https://archive.is/IRXzN) [[ 2 b]](https://web.archive.org/web/20260622051800/https://www.gdeltproject.org/data.html) | | Sentiment | Available in enrichment [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Global Content Analysis Measures covering more than 2,200 emotions and themes [[ 2 a]](https://archive.is/IRXzN) [[ 2 b]](https://web.archive.org/web/20260622051800/https://www.gdeltproject.org/data.html) | | Structured context | Source-attributed entities and event context [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | People, locations, organizations, themes, emotions, events, and counts [[ 2 a]](https://archive.is/IRXzN) [[ 2 b]](https://web.archive.org/web/20260622051800/https://www.gdeltproject.org/data.html) | | Entity and ticker mapping | Source-attributed entities [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Global Knowledge Graph entities and themes [[ 2 a]](https://archive.is/IRXzN) [[ 2 b]](https://web.archive.org/web/20260622051800/https://www.gdeltproject.org/data.html) | | Delivery and API | Six-endpoint agent API, SDKs, MCP [[ 4 a]](https://archive.is/glqrw) | Open datasets, downloads, and analysis tools [[ 1 a]](https://archive.is/vfGKI) [[ 1 b]](https://web.archive.org/web/20260707152459/https://www.gdeltproject.org/) | | Research orientation | Built for AI-agent and model workflows [[ 4 a]](https://archive.is/glqrw) | Open research into human society and the global news ecosystem [[ 1 a]](https://archive.is/vfGKI) [[ 1 b]](https://web.archive.org/web/20260707152459/https://www.gdeltproject.org/) | | Typical use | Investment, risk, and research-agent evidence [[ 4 a]](https://archive.is/glqrw) | Global-scale media, event, network, and societal analysis [[ 1 a]](https://archive.is/vfGKI) [[ 1 b]](https://web.archive.org/web/20260707152459/https://www.gdeltproject.org/) | | Pricing | See nosible.com [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) | Free and open [[ 1 a]](https://archive.is/vfGKI) [[ 1 b]](https://web.archive.org/web/20260707152459/https://www.gdeltproject.org/) | ## Potential fit by workflow GDELT's cited materials describe free and open global research, long historical coverage, and multilingual event data. [[ 1 a]](https://archive.is/vfGKI) [[ 1 b]](https://web.archive.org/web/20260707152459/https://www.gdeltproject.org/) NOSIBLE's cited materials describe dated source documents, ranked events, ticker context, and a managed workflow for agents. [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) In NOSIBLE's view, consider the product whose published operating model matches the requirement; a team may also evaluate the products together after validating each dataset. ## Common GDELT comparison questions ### How does NOSIBLE feel it differentiates itself from GDELT? NOSIBLE is an AI-native company with two products: SEARCH and WORLD. [[ 5 a]](https://archive.is/tGYc9) [[ 7 a]](https://archive.is/iDZeo) SEARCH lets agents find dated open-web sources they can cite and inspect directly. [[ 4 a]](https://archive.is/glqrw) WORLD is a live open-web event database for models and backtests, with an embedding per event. [[ 5 a]](https://archive.is/tGYc9) [[ 6 a]](https://archive.is/tnkpG) [[ 8 a]](https://archive.is/1CUvQ) NOSIBLE is committed to open-source software and makes its models publicly available on Hugging Face. [[ 9 a]](https://archive.is/7SSz0) [[ 10 a]](https://archive.is/kHxMG) Related [WORLD v1.2 trial](https://nosible.com/start-trial#data-coverage) [Sentiment model](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) [Forward-looking model](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) ### When is GDELT the more direct choice? GDELT publishes free and open global-news research infrastructure, multilingual coverage, and GDELT 1.0 event data reaching back to 1979. [[ 2 a]](https://archive.is/IRXzN) [[ 2 b]](https://web.archive.org/web/20260622051800/https://www.gdeltproject.org/data.html) NOSIBLE publishes managed source evidence, ranked events, ticker context, and agent workflows. [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) In NOSIBLE's view, GDELT is the more direct choice when its published operating model is the requirement. Related [WORLD event database](https://nosible.world/world) [Stock Screening](https://nosible.com/#proof) ### How should an investment team evaluate GDELT and NOSIBLE? Run representative issuer, theme, and event queries through both products. Evaluate source traceability, entity mapping, historical timing, duplication, language handling, and the effort required to connect results to securities before selecting either dataset for a live or historical model. Related [WORLD event database](https://nosible.world/world) [Bulk Web Search](https://docs.nosible.com/endpoints/search-bulk) ### What does GDELT publish about history and update frequency? GDELT's documentation says GDELT 1.0 includes event data back to January 1, 1979 and GDELT 2.0 updates every 15 minutes. [[ 2 a]](https://archive.is/IRXzN) [[ 2 b]](https://web.archive.org/web/20260622051800/https://www.gdeltproject.org/data.html) For any historical model, buyers should test timestamps, revisions, duplication, translation, and feature construction against their own point-in-time requirements. [[ 1 a]](https://archive.is/vfGKI) [[ 1 b]](https://web.archive.org/web/20260707152459/https://www.gdeltproject.org/) [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Related [WORLD event database](https://nosible.world/world) ### What does GDELT's language coverage include? GDELT says it monitors news in more than 100 languages and translates 65 languages in real time. [[ 2 a]](https://archive.is/IRXzN) [[ 2 b]](https://web.archive.org/web/20260622051800/https://www.gdeltproject.org/data.html) NOSIBLE publishes a 95-language long-form source corpus. [[ 4 a]](https://archive.is/glqrw) The practical comparison should test language quality, regional source mix, and retrieval precision for the markets that matter. [[ 1 a]](https://archive.is/vfGKI) [[ 1 b]](https://web.archive.org/web/20260707152459/https://www.gdeltproject.org/) [[ 3 a]](https://archive.is/r7qDv) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) Related [WORLD event database](https://nosible.world/world) ### Can GDELT and NOSIBLE be used together? Potentially. GDELT can support broad open research into global media and society, while NOSIBLE can supply dated source retrieval, ranked market events, and ticker context for research agents. [[ 1 a]](https://archive.is/vfGKI) [[ 1 b]](https://web.archive.org/web/20260707152459/https://www.gdeltproject.org/) [[ 4 a]](https://archive.is/glqrw) [[ 5 a]](https://archive.is/tGYc9) Teams should document provenance and validate each field before combining the datasets. Related [WORLD event database](https://nosible.world/world) [Bulk Web Search](https://docs.nosible.com/endpoints/search-bulk) Continue comparing ## Related Vendor Comparisons Compare GDELT with adjacent options for task execution, agent tooling, data access, and research workflows, then evaluate the output shape your application actually needs for production. [Common Crawl](https://nosible.com/compare/nosible-vs-common-crawl) [RavenPack](https://nosible.com/compare/nosible-vs-ravenpack) [News API](https://nosible.com/compare/nosible-vs-news-api) [MKT MediaStats](https://nosible.com/compare/nosible-vs-mkt-mediastats) Dive deeper ## Take your next step today Review the delivered field definitions, classification boundaries, and example values before comparing vendor workflows, so your team can assess what each product actually returns to downstream systems. [Data dictionaries](https://nosible.com/data-dictionaries) [Ontology reference](https://nosible.com/ontologies) [API reference](https://docs.nosible.com/) [Start trial](https://nosible.com/start-trial) Sources reviewed July 14, 2026 : [[ 1 ] GDELT Project overview](https://www.gdeltproject.org/) ( [dated snapshot](https://archive.is/vfGKI); [Wayback copy](https://web.archive.org/web/20260707152459/https://www.gdeltproject.org/) ) , [[ 2 ] GDELT data documentation](https://www.gdeltproject.org/data.html) ( [dated snapshot](https://archive.is/IRXzN); [Wayback copy](https://web.archive.org/web/20260622051800/https://www.gdeltproject.org/data.html) ) , [[ 3 ] NOSIBLE product overview](https://nosible.com/) ( [dated snapshot](https://archive.is/r7qDv); [Wayback copy](https://web.archive.org/web/20260713185114/https://nosible.com/) ) , [[ 4 ] NOSIBLE SEARCH](https://docs.nosible.com/) ( [dated snapshot](https://archive.is/glqrw) ) , [[ 5 ] NOSIBLE WORLD](https://nosible.world/world) ( [dated snapshot](https://archive.is/tGYc9) ) , [[ 6 ] NOSIBLE WORLD v1.2 trial and coverage](https://nosible.com/start-trial) ( [dated snapshot](https://archive.is/tnkpG) ) , [[ 7 ] NOSIBLE AI-native research overview](https://nosible.com/blog) ( [dated snapshot](https://archive.is/iDZeo) ) , [[ 8 ] NOSIBLE embedding-based research](https://nosible.com/blog/an-embedding-based-approach-to-trade-and-economic-policy-uncertainty) ( [dated snapshot](https://archive.is/1CUvQ) ) , [[ 9 ] NOSIBLE Financial Sentiment v1.2 Base](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) ( [dated snapshot](https://archive.is/7SSz0) ) , [[ 10 ] NOSIBLE Forward-Looking v1.2 Base](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) ( [dated snapshot](https://archive.is/kHxMG) ) . This page compares NOSIBLE WORLD with the GDELT Project. GDELT and related names are marks of their respective owner. NOSIBLE is not affiliated with, sponsored by, or endorsed by the GDELT Project. Comparison standards, legal context & corrections This comparison was prepared by NOSIBLE, which has a commercial interest in the products being compared. It is based on the cited public materials as they appeared on July 14, 2026. NOSIBLE has not tested every competitor feature, and the page is not a complete statement of either product. Products change; confirm current requirements, availability, and commercial terms with each vendor. Factual claims are attributed to cited first-party materials; evaluative statements reflect NOSIBLE's opinion. Each source includes a dated archive.is snapshot and, where the Internet Archive captured that URL, a timestamp-specific Wayback copy. Archive availability is controlled by those services. Competitor names and marks are used only to identify the products being compared. No affiliation, sponsorship, or endorsement is implied. The FTC says truthful, non-deceptive comparative advertising may identify competitors ( [policy](https://www.ftc.gov/legal-library/browse/statement-policy-regarding-comparative-advertising); [dated snapshot](https://archive.is/nZXY4); [Wayback copy](https://web.archive.org/web/20260215042003/https://www.ftc.gov/legal-library/browse/statement-policy-regarding-comparative-advertising) ). For one U.S. example of nominative-use analysis, see *New Kids on the Block v. News America Publishing, Inc.*, 971 F.2d 302, 308 (9th Cir. 1992) ( [opinion](https://law.justia.com/cases/federal/appellate-courts/F2/971/302/72076/); [dated snapshot](https://archive.is/dnk8j); [Wayback copy](https://web.archive.org/web/20250211213616/https://law.justia.com/cases/federal/appellate-courts/F2/971/302/72076/) ). If you represent GDELT and believe a factual statement is inaccurate, email [stuart@nosible.com](mailto:stuart@nosible.com?subject=Correction%20request%3A%20NOSIBLE%20WORLD%20vs%20GDELT) with the specific claim and a supporting first-party URL. NOSIBLE will review and correct substantiated errors. > How NOSIBLE WORLD, a web-scale search and market-event intelligence engine for AI agents, compares with GDELT, a free and open global-news research platform. Coverage, history, delivery, and use cases. **URL:** https://nosible.com/compare/nosible-vs-gdelt --- --- title: "NOSIBLE vs Common Crawl" description: "Compare NOSIBLE and Common Crawl for open web data, crawl archives, indexes, source evidence, event history, processing effort, and research workflows." url: "https://nosible.com/compare/nosible-vs-common-crawl" --- Comparison / Reviewed July 25, 2026 # NOSIBLE vs Common Crawl Common Crawl describes free web-crawl infrastructure for teams prepared to process it. [[ 2 a]](https://archive.is/t1QK6) NOSIBLE is the managed research layer for agents that need dated sources and ranked events without beginning at raw crawl files. [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) NOSIBLE-AUTHORED COMPARISON · NOSIBLE HAS A COMMERCIAL INTEREST IN THIS COMPARISON · FACTS ATTRIBUTED TO FIRST-PARTY VENDOR MATERIALS · EVALUATIVE STATEMENTS ARE NOSIBLE'S OPINION · REVIEWED JULY 25, 2026 · DUAL PUBLIC ARCHIVES WHERE SUPPORTED · [STANDARDS & CORRECTIONS](https://nosible.com/compare/nosible-vs-common-crawl#comparison-standards). IF YOU REPRESENT COMMON CRAWL AND BELIEVE A FACTUAL STATEMENT IS INACCURATE, EMAIL [STUART@NOSIBLE.COM](mailto:stuart@nosible.com?subject=Correction%20request%3A%20NOSIBLE%20vs%20Common%20Crawl) WITH THE SPECIFIC CLAIM AND A SUPPORTING FIRST-PARTY URL. NOSIBLE WILL REVIEW AND CORRECT SUBSTANTIATED ERRORS. - Common Crawl describes a petabyte-scale web corpus collected since 2008. [[ 1 a]](https://archive.is/BDcUj) - Common Crawl publishes raw pages, metadata, text extracts, and a URL index through its crawl data. [[ 1 a]](https://archive.is/BDcUj) - Common Crawl says its data is free and available through AWS Open Data and direct download methods. [[ 1 a]](https://archive.is/BDcUj) [[ 2 a]](https://archive.is/t1QK6) - NOSIBLE SEARCH and WORLD provide managed dated source retrieval and ranked events, including an embedding per event, rather than raw web-crawl files. [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) [[ 8 a]](https://archive.is/1CUvQ) - In NOSIBLE's view, NOSIBLE is the stronger fit when a team needs ready-to-use agent evidence instead of operating its own web-data pipeline. ## The raw corpus route versus a usable research layer Common Crawl publishes a large open corpus of raw pages, metadata, text extracts, and a URL index. [[ 1 a]](https://archive.is/BDcUj) NOSIBLE starts after that infrastructure decision, with managed dated sources and ranked events for research agents. [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) In NOSIBLE's view, the honest comparison is not free versus paid; it is whether a team wants to own the pipeline from WARC processing onward or begin with evidence and event data. ## Open-web infrastructure and managed research evidence The competitor column summarizes what the named vendor's cited first-party materials describe; the NOSIBLE column summarizes NOSIBLE's cited materials. Primary use NOSIBLE Managed dated source and event intelligence [[ 4 a]](https://archive.is/YwGWR) Common Crawl Open web-crawl corpus and index for data processing [[ 2 a]](https://archive.is/t1QK6) Corpus NOSIBLE Open-web source retrieval within NOSIBLE products [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 4 a]](https://archive.is/YwGWR) Common Crawl Petabyte-scale crawl corpus collected since 2008 [[ 1 a]](https://archive.is/BDcUj) Data forms NOSIBLE Source evidence and ranked events [[ 5 a]](https://archive.is/fv0cj) Common Crawl Raw pages, metadata, text extracts, and URL index [[ 1 a]](https://archive.is/BDcUj) Access NOSIBLE API, SDKs, MCP, SEARCH, and WORLD [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) Common Crawl Free access through AWS Open Data and downloads [[ 1 a]](https://archive.is/BDcUj) [[ 2 a]](https://archive.is/t1QK6) Processing model NOSIBLE Managed retrieval and enrichment Common Crawl User processes WARC and related crawl data [[ 2 a]](https://archive.is/t1QK6) Point-in-time method NOSIBLE Documented buyer evaluation Common Crawl Crawl snapshots and index records; validate dates for the task Event representation NOSIBLE Ranked dated WORLD events [[ 5 a]](https://archive.is/fv0cj) Common Crawl Crawl corpus rather than a published ranked event database Operational effort NOSIBLE Managed product workflows Common Crawl Evaluate the processing, filtering, and governance effort for the intended use Pricing NOSIBLE Confirm current terms with NOSIBLE Common Crawl Common Crawl publishes free public data; evaluate processing requirements [[ 2 a]](https://archive.is/t1QK6) | Dimension | NOSIBLE | Common Crawl | | --- | --- | --- | | Primary use | Managed dated source and event intelligence [[ 4 a]](https://archive.is/YwGWR) | Open web-crawl corpus and index for data processing [[ 2 a]](https://archive.is/t1QK6) | | Corpus | Open-web source retrieval within NOSIBLE products [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 4 a]](https://archive.is/YwGWR) | Petabyte-scale crawl corpus collected since 2008 [[ 1 a]](https://archive.is/BDcUj) | | Data forms | Source evidence and ranked events [[ 5 a]](https://archive.is/fv0cj) | Raw pages, metadata, text extracts, and URL index [[ 1 a]](https://archive.is/BDcUj) | | Access | API, SDKs, MCP, SEARCH, and WORLD [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) | Free access through AWS Open Data and downloads [[ 1 a]](https://archive.is/BDcUj) [[ 2 a]](https://archive.is/t1QK6) | | Processing model | Managed retrieval and enrichment | User processes WARC and related crawl data [[ 2 a]](https://archive.is/t1QK6) | | Point-in-time method | Documented buyer evaluation | Crawl snapshots and index records; validate dates for the task | | Event representation | Ranked dated WORLD events [[ 5 a]](https://archive.is/fv0cj) | Crawl corpus rather than a published ranked event database | | Operational effort | Managed product workflows | Evaluate the processing, filtering, and governance effort for the intended use | | Pricing | Confirm current terms with NOSIBLE | Common Crawl publishes free public data; evaluate processing requirements [[ 2 a]](https://archive.is/t1QK6) | ## Pipeline ownership versus time to usable evidence In NOSIBLE's view, Common Crawl may be the more direct choice for a team that needs free crawl files and the freedom to operate its own web-data pipeline. NOSIBLE is designed for managed dated sources and ranked events, with an embedding per event for downstream research. [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) [[ 8 a]](https://archive.is/1CUvQ) In NOSIBLE's view, a custom corpus program can still use NOSIBLE when it needs a ready-to-use evidence or event layer. ## Common Crawl comparison questions ### How does NOSIBLE feel it differentiates itself from Common Crawl? NOSIBLE is an AI-native company with two products: SEARCH and WORLD. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 5 a]](https://archive.is/fv0cj) [[ 7 a]](https://archive.is/iDZeo) SEARCH lets agents find dated open-web sources they can cite and inspect directly. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 4 a]](https://archive.is/YwGWR) WORLD is a live open-web event database for models and backtests, with an embedding per event. [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 5 a]](https://archive.is/fv0cj) [[ 6 a]](https://archive.is/tnkpG) [[ 8 a]](https://archive.is/1CUvQ) NOSIBLE is committed to open-source software and makes its models publicly available on Hugging Face. [[ 9 a]](https://archive.is/7SSz0) [[ 10 a]](https://archive.is/kHxMG) Related [WORLD v1.2 trial](https://nosible.com/start-trial#data-coverage) [Sentiment model](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) [Forward-looking model](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) ### What does Common Crawl publish? Common Crawl describes a petabyte-scale web corpus collected since 2008, with raw pages, metadata, text extracts, and a URL index. [[ 1 a]](https://archive.is/BDcUj) It says the data is free through AWS Open Data and download methods. [[ 1 a]](https://archive.is/BDcUj) [[ 2 a]](https://archive.is/t1QK6) In NOSIBLE's view, Common Crawl may be the more direct fit for teams building their own web-data pipeline. Related [WORLD event database](https://nosible.world/world) [HTML to JSON](https://docs.nosible.com/endpoints/search-scrape) ### How does NOSIBLE differ from Common Crawl? Common Crawl provides crawl data for users to process, while NOSIBLE provides managed dated source retrieval and ranked events for agents and research. [[ 2 a]](https://archive.is/t1QK6) [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) In NOSIBLE's view, the comparison turns on operational responsibility: a buyer should assess acquisition, storage, filtering, extraction, date handling, provenance, and governance before choosing a data foundation. Related [WORLD event database](https://nosible.world/world) [HTML to JSON](https://docs.nosible.com/endpoints/search-scrape) ### Does Common Crawl provide a comparable event database? The cited Common Crawl materials describe a web corpus, crawl files, and a URL index, not a published ranked event database comparable to NOSIBLE WORLD. [[ 1 a]](https://archive.is/BDcUj) [[ 5 a]](https://archive.is/fv0cj) This page therefore does not compare aggregate event counts. In NOSIBLE's view, a crawl record and a market-event record should be evaluated as different data products. Related [WORLD event database](https://nosible.world/world) [HTML to JSON](https://docs.nosible.com/endpoints/search-scrape) ### What processing work should a Common Crawl evaluation include? A serious evaluation should include downloading or processing WARC data, URL-index selection, deduplication, language and quality filters, text extraction, date interpretation, storage, and governance. [[ 1 a]](https://archive.is/BDcUj) [[ 2 a]](https://archive.is/t1QK6) These steps determine whether a crawl corpus is usable for the specific application and should be measured before comparing it with a managed evidence product. Related [WORLD event database](https://nosible.world/world) [HTML to JSON](https://docs.nosible.com/endpoints/search-scrape) ### Can Common Crawl and NOSIBLE be used together? Potentially. Common Crawl can support a custom open-web corpus program, while NOSIBLE can supply managed dated source retrieval and ranked event context. [[ 1 a]](https://archive.is/BDcUj) [[ 3 a]](https://archive.is/8abI7) [[ 3 b]](https://web.archive.org/web/20260713185114/https://nosible.com/) [[ 4 a]](https://archive.is/YwGWR) [[ 5 a]](https://archive.is/fv0cj) In NOSIBLE's view, teams should preserve crawl identifiers, source links, processing rules, timestamps, and dataset lineage before combining the two in research or model workflows. Related [WORLD event database](https://nosible.world/world) [Bulk Web Search](https://docs.nosible.com/endpoints/search-bulk) [HTML to JSON](https://docs.nosible.com/endpoints/search-scrape) Continue comparing ## Related Vendor Comparisons Compare Common Crawl with adjacent options for task execution, agent tooling, data access, and research workflows, then evaluate the output shape your application actually needs for production. [GDELT](https://nosible.com/compare/nosible-vs-gdelt) [Exa](https://nosible.com/compare/nosible-vs-exa) [Tavily](https://nosible.com/compare/nosible-vs-tavily) [Parallel](https://nosible.com/compare/nosible-vs-parallel) Dive deeper ## Take your next step today Review the delivered field definitions, classification boundaries, and example values before comparing vendor workflows, so your team can assess what each product actually returns to downstream systems. [Data dictionaries](https://nosible.com/data-dictionaries) [Ontology reference](https://nosible.com/ontologies) [API reference](https://docs.nosible.com/) [Start trial](https://nosible.com/start-trial) Sources reviewed July 25, 2026 : [[ 1 ] Common Crawl overview](https://commoncrawl.org/overview) ( [dated snapshot](https://archive.is/BDcUj) ) , [[ 2 ] Common Crawl get started](https://commoncrawl.org/get-started) ( [dated snapshot](https://archive.is/t1QK6) ) , [[ 3 ] NOSIBLE product overview](https://nosible.com/) ( [dated snapshot](https://archive.is/8abI7); [Wayback copy](https://web.archive.org/web/20260713185114/https://nosible.com/) ) , [[ 4 ] NOSIBLE SEARCH](https://docs.nosible.com/) ( [dated snapshot](https://archive.is/YwGWR) ) , [[ 5 ] NOSIBLE WORLD](https://nosible.world/world) ( [dated snapshot](https://archive.is/fv0cj) ) , [[ 6 ] NOSIBLE WORLD v1.2 trial and coverage](https://nosible.com/start-trial) ( [dated snapshot](https://archive.is/tnkpG) ) , [[ 7 ] NOSIBLE AI-native research overview](https://nosible.com/blog) ( [dated snapshot](https://archive.is/iDZeo) ) , [[ 8 ] NOSIBLE embedding-based research](https://nosible.com/blog/an-embedding-based-approach-to-trade-and-economic-policy-uncertainty) ( [dated snapshot](https://archive.is/1CUvQ) ) , [[ 9 ] NOSIBLE Financial Sentiment v1.2 Base](https://huggingface.co/NOSIBLE/financial-sentiment-v1.2-base) ( [dated snapshot](https://archive.is/7SSz0) ) , [[ 10 ] NOSIBLE Forward-Looking v1.2 Base](https://huggingface.co/NOSIBLE/forward-looking-v1.2-base) ( [dated snapshot](https://archive.is/kHxMG) ) . Common Crawl is used solely to identify the compared product. NOSIBLE is not affiliated with, sponsored by, or endorsed by Common Crawl. Comparison standards, legal context & corrections This comparison was prepared by NOSIBLE, which has a commercial interest in the products being compared. It is based on the cited public materials as they appeared on July 25, 2026. NOSIBLE has not tested every competitor feature, and the page is not a complete statement of either product. Products change; confirm current requirements, availability, and commercial terms with each vendor. Factual claims are attributed to cited first-party materials; evaluative statements reflect NOSIBLE's opinion. Each source includes a dated archive.is snapshot and, where the Internet Archive captured that URL, a timestamp-specific Wayback copy. Archive availability is controlled by those services. Competitor names and marks are used only to identify the products being compared. No affiliation, sponsorship, or endorsement is implied. The FTC says truthful, non-deceptive comparative advertising may identify competitors ( [policy](https://www.ftc.gov/legal-library/browse/statement-policy-regarding-comparative-advertising); [dated snapshot](https://archive.is/nZXY4); [Wayback copy](https://web.archive.org/web/20260215042003/https://www.ftc.gov/legal-library/browse/statement-policy-regarding-comparative-advertising) ). For one U.S. example of nominative-use analysis, see *New Kids on the Block v. News America Publishing, Inc.*, 971 F.2d 302, 308 (9th Cir. 1992) ( [opinion](https://law.justia.com/cases/federal/appellate-courts/F2/971/302/72076/); [dated snapshot](https://archive.is/dnk8j); [Wayback copy](https://web.archive.org/web/20250211213616/https://law.justia.com/cases/federal/appellate-courts/F2/971/302/72076/) ). If you represent Common Crawl and believe a factual statement is inaccurate, email [stuart@nosible.com](mailto:stuart@nosible.com?subject=Correction%20request%3A%20NOSIBLE%20vs%20Common%20Crawl) with the specific claim and a supporting first-party URL. NOSIBLE will review and correct substantiated errors. > Compare NOSIBLE and Common Crawl for open web data, crawl archives, indexes, source evidence, event history, processing effort, and research workflows. **URL:** https://nosible.com/compare/nosible-vs-common-crawl