indexer/README.md
2026-09-28 15:36:42 +02:00

1.7 KiB
Raw Blame History

indexer

Federated Forge Discovery Indexer

Builds an index of public open source software projects from configured forges.

Scrape behavior

For each configured forge, the indexer:

  1. Checks robots.txt (GET /robots.txt) before any API call. If the instance disallows /api/, the instance is skipped.
  2. Lists public source repositories via the repository search API (GET /api/v1/repos/search), paginated. Forks and mirrors are excluded.
  3. For each repository, enriches it via the Forgejo API: reads the root directory once, then fetches topics, languages, latest commit, latest release, the README, and detects licenses from the repo's files.

Errors are non-fatal: an instance or a single repo that fails is skipped and the run continues.

Request volume

Per instance:

Request Count
GET /robots.txt 1
GET /api/v1/repos/search (page) ceil(N/50) where N = total repos

Per repository:

Request Endpoint Purpose
list root contents GET /api/v1/repos/{o}/{r}/contents/ shared: README selection + license scan (1 call)
languages GET /api/v1/repos/{o}/{r}/languages Project.Languages
topics GET /api/v1/repos/{o}/{r}/topics Project.Topics
latest commit GET /api/v1/repos/{o}/{r}/commits?limit=1 Project.LatestCommit
latest release GET /api/v1/repos/{o}/{r}/releases/latest Project.LatestRelease
README GET /api/v1/repos/{o}/{r}/raw/{branch}/{file} Project.Readme
license files 1+ × GET /api/v1/repos/{o}/{r}/raw/{branch}/{file} SPDX detection

So roughly 6–8 API requests per repository plus the paginated search requests and a single robots.txt check per instance.