44 lines
1.7 KiB
Markdown
44 lines
1.7 KiB
Markdown
# indexer
|
||
|
||
Federated Forge Discovery Indexer
|
||
|
||
Builds an index of public open source software projects from configured forges.
|
||
|
||
## Scrape behavior
|
||
|
||
For each configured forge, the indexer:
|
||
|
||
1. **Checks robots.txt** (`GET /robots.txt`) before any API call. If the
|
||
instance disallows `/api/`, the instance is skipped.
|
||
2. **Lists public source repositories** via the repository search API
|
||
(`GET /api/v1/repos/search`), paginated. Forks and mirrors are excluded.
|
||
3. For each repository, enriches it via the Forgejo API: reads the root
|
||
directory once, then fetches topics, languages, latest commit, latest
|
||
release, the README, and detects licenses from the repo's files.
|
||
|
||
Errors are non-fatal: an instance or a single repo that fails is skipped and
|
||
the run continues.
|
||
|
||
### Request volume
|
||
|
||
**Per instance:**
|
||
|
||
| Request | Count |
|
||
|---|---|
|
||
| `GET /robots.txt` | 1 |
|
||
| `GET /api/v1/repos/search` (page) | `ceil(N/50)` where N = total repos |
|
||
|
||
**Per repository:**
|
||
|
||
| Request | Endpoint | Purpose |
|
||
|---|---|---|
|
||
| list root contents | `GET /api/v1/repos/{o}/{r}/contents/` | shared: README selection + license scan (1 call) |
|
||
| languages | `GET /api/v1/repos/{o}/{r}/languages` | `Project.Languages` |
|
||
| topics | `GET /api/v1/repos/{o}/{r}/topics` | `Project.Topics` |
|
||
| latest commit | `GET /api/v1/repos/{o}/{r}/commits?limit=1` | `Project.LatestCommit` |
|
||
| latest release | `GET /api/v1/repos/{o}/{r}/releases/latest` | `Project.LatestRelease` |
|
||
| README | `GET /api/v1/repos/{o}/{r}/raw/{branch}/{file}` | `Project.Readme` |
|
||
| license files | 1+ × `GET /api/v1/repos/{o}/{r}/raw/{branch}/{file}` | SPDX detection |
|
||
|
||
So roughly **6–8 API requests per repository** plus the paginated search
|
||
requests and a single robots.txt check per instance.
|