Skip to content

CapTech IDF — IN PROGRESS

Maps 79,099 Île-de-France companies from public Insee data across 16 technology domains. A company joins a domain only with a cited, dated source; 23 have one so far, in AI, cybersecurity and embedded systems. The code is private.

TL;DR
  • 79,099 companies
  • 1,343 indexed segments

My partThe data pipeline from the Sirene stocks (Parquet, DuckDB) into Supabase.

ROLE
Solo
CONTEXT
Personal project
STATUS
In progress
RESULT
Corpus cleaning cut indexed segments from 2,105 to 1,343; 45 of 46 organisations kept business text. Retrieval is benchmarked on 18 queries as positive–unlabelled, so no accuracy or F1 is claimed.
LAST UPDATED
26 Sep 2026
CapTech IDF scene for AI companies in Paris: each evaluated company sits on a ring set by the strength of the evidence found for the domain.Live site · AI companies in Paris
1

Data

Insee Sirene (company and establishment stocks) and the public company-search API, both under Licence Ouverte; domain names from the INPI company register and text from company websites, used offline only.

2

What I built

Solo

  1. The data pipeline from the Sirene stocks (Parquet, DuckDB) into Supabase.
  2. Offline enrichment: company websites found through INPI domain names, crawled within robots.txt at most once every 1.5 seconds per host, and an evidence corpus filtered by URL, page type and text block.
  3. The interactive map (MapLibre) and the 2D and 3D scenes of companies and links (ECharts, three.js).
  4. More than 1,000 automated tests.
Sirene stocks
Parquet · DuckDB
load
Supabase
crawl sites
robots.txt · 1 req/s/host
evidence corpus
URL · page · block filters
cited domain
source · date
map · 2D/3D scene
MapLibre · three.js
tests — 1,000+ automated tests (Vitest)
Fig. 1 — Data flow as documented in the private repo. Outlined in accent: the evidence path; a company joins a domain only with a cited, dated source.
3

Key choices

Evidence before labels
A company is attached to a domain only with a cited, dated source; companies matched by activity code alone stay candidates.
Similarity finds candidates, it is never shown
The interface shows an explainable proximity band backed by citations, not a score out of 100.
No personal data
Directors and beneficial owners are excluded at the request level.
4

Results

Corpus cleaning cut indexed segments from 2,105 to 1,343; 45 of 46 organisations kept business text. Retrieval is benchmarked on 18 queries as positive–unlabelled, so no accuracy or F1 is claimed.

5

Limits

  • Activity codes alone over-include: about 40,000 AI candidates in Paris and 61,216 in Île-de-France before any evidence.
  • The scene draws at most 400 objects; companies beyond that are grouped by activity code.
  • Proximity is shown as bands derived from evidence; no numeric weights are set, because 29 reference positives are too few to calibrate them.

Questions about this project? → Email me