Federated Forge Discovery Indexer
- Go 87.4%
- Nix 12.6%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
| scraper | ||
| store | ||
| .editorconfig | ||
| .gitignore | ||
| config.go | ||
| flake.lock | ||
| flake.nix | ||
| go.mod | ||
| go.sum | ||
| gomod2nix.toml | ||
| LICENSE | ||
| main.go | ||
| option.nix | ||
| README.md | ||
indexer
Federated Forge Discovery Indexer
Builds an index of public open source software projects from configured forges.
Scrape behavior
For each configured forge, the indexer:
- Checks robots.txt (
GET /robots.txt) before any API call. If the instance disallows/api/, the instance is skipped. - Lists public source repositories via the repository search API
(
GET /api/v1/repos/search), paginated. Forks and mirrors are excluded. - For each repository, enriches it via the Forgejo API: reads the root directory once, then fetches topics, languages, latest commit, latest release, the README, and detects licenses from the repo's files.
Errors are non-fatal: an instance or a single repo that fails is skipped and the run continues.
Request volume
Per instance:
| Request | Count |
|---|---|
GET /robots.txt |
1 |
GET /api/v1/repos/search (page) |
ceil(N/50) where N = total repos |
Per repository:
| Request | Endpoint | Purpose |
|---|---|---|
| list root contents | GET /api/v1/repos/{o}/{r}/contents/ |
shared: README selection + license scan (1 call) |
| languages | GET /api/v1/repos/{o}/{r}/languages |
Project.Languages |
| topics | GET /api/v1/repos/{o}/{r}/topics |
Project.Topics |
| latest commit | GET /api/v1/repos/{o}/{r}/commits?limit=1 |
Project.LatestCommit |
| latest release | GET /api/v1/repos/{o}/{r}/releases/latest |
Project.LatestRelease |
| README | GET /api/v1/repos/{o}/{r}/raw/{branch}/{file} |
Project.Readme |
| license files | 1+ × GET /api/v1/repos/{o}/{r}/raw/{branch}/{file} |
SPDX detection |
So roughly 6–8 API requests per repository plus the paginated search requests and a single robots.txt check per instance.