Files
ExploreDNS/README.md
T
Gary HansenandClaude Fable 5 c9a963ecdd
CI / test (pull_request) Successful in 14m14s
CI / docker (pull_request) Has been skipped
chore(receiver): packaging, k8s manifests, CI/release, docs
Dockerfile.receiver (CGO-free, /data volume), receiver image in CI and
tag releases, receiver binary in release archives, make build-receiver,
example k8s manifests (deployment/service/ingress/secret/pvc) under
deploy/k8s/receiver/, and README coverage including sender/receiver
token pairing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 02:43:51 +10:00

599 lines
23 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# ExploreDNS
ExploreDNS is a comprehensive DNS traversal tool that explores every possible
resolution path for a domain — from the root servers all the way down — just
like a real iterative resolver, but without stopping at the first answer. It
follows every referral exhaustively, collates all results, and presents them in
a structured, human-readable (or JSON) report.
Inspired by the classic [dnstraverse](https://github.com/squish/dnstraverse)
Ruby tool, ExploreDNS is a modern Go rewrite that produces a self-contained
binary with no runtime dependencies.
---
## Features
- **Full iterative traversal** — queries root servers and follows every referral
branch, mirroring real resolver behaviour
- **All root servers** — optionally query all 13 root server sets in parallel
- **No-glue resolution** — automatically resolves nameserver addresses when
referrals lack glue records
- **DNS server fingerprinting** — identifies server software via `version.bind`
CHAOS queries
- **Multiple output formats** — coloured text tree and machine-readable JSON
- **Configurable transport** — UDP/TCP, EDNS0 buffer size, retries, timeouts
- **Fast mode** — shares glue across branches for speed; disable for independent
paths
- **CNAME tracking** — follows CNAME chains and detects loops
- **Web interface** — browser-based UI backed by an HTTP API server with
real-time Server-Sent Events progress streaming
---
## Installation
### Pre-built binary
Download the latest release binary for your platform from the
[Releases page](https://gitea.hansenits.com.au/hits/ExploreDNS/releases).
### Build from source
Requires Go 1.21 or later.
```sh
git clone https://gitea.hansenits.com.au/hits/ExploreDNS.git
cd ExploreDNS
make build # produces bin/exploredns
make build-server # produces bin/exploredns-server
make build-all # produces both binaries
```
### go install
```sh
go install gitea.hansenits.com.au/hits/ExploreDNS/cmd/exploredns@latest
go install gitea.hansenits.com.au/hits/ExploreDNS/cmd/server@latest
```
---
## Quick Start
```sh
# Basic A record traversal
exploredns www.example.com
# MX records for a domain
exploredns --type MX example.com
# Use all 13 root server sets
exploredns --all-root-servers www.example.com
# JSON output
exploredns --json www.example.com
# Quiet (results only)
exploredns --quiet www.example.com
# Debug mode
exploredns --debug www.example.com
# Debug plus library-level diagnostics
exploredns --dd www.example.com
# Force TCP
exploredns --allow-tcp --always-tcp www.example.com
# Disable fast mode (independent paths per branch)
exploredns --fast=false www.example.com
# Increase traversal depth limit
exploredns --max-depth 30 www.example.com
```
---
## CLI Reference
```
Usage:
exploredns [flags] <domain>
Query Options:
--type <TYPE> Record type to query (default: a)
Supported: A, AAAA, NS, CNAME, MX, TXT, SOA, PTR, ANY
--root-server <HOST> Initial root server, hostname or IP literal
(default: ask the upstream resolver for one root)
--all-root-servers Traverse from all root servers (default: false)
--root-aaaa Include IPv6 root addresses (not implemented yet)
--follow-aaaa Only follow AAAA for referrals (not implemented yet)
--dns-upstream <ADDR> Upstream resolver (host:port) for root discovery
(default: system resolver)
Transport Options:
--udp-size <N> EDNS0 UDP buffer size, 512–4096; 512 turns EDNS0 off
(default: 2048)
--allow-tcp Fall back to TCP on truncation (default: true)
--always-tcp Always use TCP (requires --allow-tcp)
--retries <N> Number of 2s retries before timing out, 0–10 (default: 2)
Traversal Options:
--max-depth <N> Maximum referral depth, 1–100 (default: 20)
--fast / --fast=false Fast mode; turn off to be more accurate (default: true)
Output Options:
--json Emit a single JSON document instead of text
--verbose, -v Verbose progress ([qname] and <bailiwick> shown)
-d, --debug Print debug diagnostics to stderr
-dd Like -d plus library-level debug
--quiet, -q Suppress the header block
--show-progress Show traversal progress (default: true)
--show-resolves Show glue-resolution progress (default: false)
--show-servers Show servers encountered (default: false)
--show-versions Show server version fingerprints (default: true)
--show-all-stats Show statistics after every node (default: false)
--show-results Show the results (default: true)
--show-summary-results Show the summary results (default: true)
General Options:
--version, -V Print version ("exploredns <version>") and exit
```
Every `--show-X` flag can be negated with `--show-X=false` or `--no-show-X`.
---
## Web Interface
ExploreDNS ships a second binary — `exploredns-server` — that exposes a
browser-based UI and a JSON REST API backed by the same traversal engine as
the CLI.
### Starting the server
```sh
# Default: listen on :8080
./bin/exploredns-server
# Custom address
./bin/exploredns-server --addr :9090
./bin/exploredns-server --addr 127.0.0.1:8080
```
Or via Make:
```sh
make build-server
./bin/exploredns-server
```
Open `http://localhost:8080` in your browser. The SPA lets you enter a domain,
choose a record type, and watch the traversal progress in real time as a live
detail tree modelled on the dns.squish.net detail page: one node per referral,
indented per depth, with glue-resolution subtrees collapsed behind per-node
"show resolve" toggles (a "raw log" toggle reveals the flat event feed for
debugging). When the traversal completes the full result list is displayed,
followed by a Servers section: every nameserver queried during the traversal
is fingerprinted (`version.bind`) and shown on an OpenStreetMap/Leaflet map
plus a Country / City / Servers / Software guess table. Geolocation happens
client-side in your browser via the free [geojs.io](https://www.geojs.io/)
API (`get.geojs.io`); servers that cannot be located are still listed with a
dash location, and the table works without the map when offline.
### API endpoints
| Method | Path | Description |
|--------|------|-------------|
| `POST` | `/api/traverse` | Start an asynchronous traversal |
| `GET` | `/api/traverse/{id}` | Poll traversal status and results |
| `GET` | `/api/traverse/{id}/stream` | Server-Sent Events live progress stream |
| `GET` | `/api/traverse/{id}/servers` | Fingerprinted list of every server queried |
| `GET` | `/api/health` | Health check — returns `{"status":"ok","version":"<build version>"}` |
#### POST /api/traverse
Request body (JSON):
```json
{
"domain": "www.example.com",
"type": "A",
"all_roots": false
}
```
`type` defaults to `"A"` if omitted. `all_roots` queries all 13 root server
sets in parallel (equivalent to `--all-root-servers` in the CLI).
Response (`202 Accepted`):
```json
{
"id": "3fa85f64-5717-4562-b3fc-2c963f66afa6",
"status": "running"
}
```
#### GET /api/traverse/{id}
Returns a snapshot of the job including the full result list once complete.
`status` is one of `running`, `complete`, or `error`. Once post-traversal
fingerprinting has finished the snapshot also carries a `servers` array (the
same list served by `GET /api/traverse/{id}/servers`; omitted before then).
```json
{
"id": "3fa85f64-5717-4562-b3fc-2c963f66afa6",
"status": "complete",
"domain": "www.example.com",
"query_type": "A",
"started_at": "2024-01-01T12:00:00Z",
"done_at": "2024-01-01T12:00:02Z",
"results": [
{
"depth": 2,
"probability": 1.0,
"response_type": "Answer",
"server": "192.0.2.53:53",
"answers": ["www.example.com. 3600 IN A 93.184.216.34"]
}
]
}
```
#### GET /api/traverse/{id}/servers
Every `(server name, IP)` pair queried during the traversal — including
glue-resolution subtree servers — fingerprinted via a `version.bind` CHAOS
probe once the traversal reaches a terminal state. Fingerprinting never
delays the traversal results: while it (or the traversal itself) is still in
flight the endpoint answers `202 Accepted` with `{"status":"pending"}`.
Unknown ids answer `404`. When ready:
```json
{
"status": "complete",
"servers": [
{"name": "a.iana-servers.net", "ip": "199.43.135.53", "version": ""},
{"name": "l.gtld-servers.net", "ip": "192.41.162.30", "version": "..."}
]
}
```
`version` is `""` for servers that don't answer the probe.
#### GET /api/traverse/{id}/stream
An [SSE](https://developer.mozilla.org/en-US/docs/Web/API/Server-sent_events)
stream of `ProgressEvent` objects, one per `data:` message. Past events
recorded before the client connected are replayed immediately, then live events
follow. Two synthetic stages bracket the end of a job: `{"stage":"complete"}`
when the traversal reaches its terminal status (results are fetchable) and
`{"stage":"servers"}` when the fingerprinted server list is ready. The stream
ends with `event: done`.
```
data: {"stage":"start","depth":1,"name":"www.example.com","qtype":"A","bailiwick":"com"}
data: {"stage":"complete","depth":1,"name":"www.example.com","qtype":"A","server":"192.0.2.53:53","bailiwick":"com"}
event: done
data: {}
```
Completed jobs are kept in memory for one hour before being purged.
### Server configuration
The server is safe to expose publicly by default and reads these environment
variables at startup:
| Variable | Default | Meaning |
|---|---|---|
| `EXPLOREDNS_JOB_TIMEOUT` | `5m` | Hard deadline per traversal (Go duration). Timed-out jobs report `error` with any partial results. |
| `EXPLOREDNS_MAX_JOBS` | `8` | Maximum concurrent traversals; further `POST /api/traverse` requests get `429`. |
| `EXPLOREDNS_CORS_ORIGIN` | *(unset)* | Off by default (the SPA is same-origin). Set an origin — or `*` for development — to enable cross-origin API access. |
| `EXPLOREDNS_RATE_LIMIT` | `30/1h` | Per-client-IP token-bucket limit on `POST /api/traverse` in `N/duration` form (e.g. `10/10m`); invalid values fall back to the default. Over-limit requests get `429`. Buckets refill continuously. Direct localhost connections are exempt (dev loop, tests), but proxied requests are always limited by the real client IP from `Fly-Client-IP` / `X-Forwarded-For`. |
| `EXPLOREDNS_WEBHOOK_URL` | *(unset)* | Off by default. When set, the server POSTs a usage-reporting JSON event to this URL on every traversal start and completion (see below). |
| `EXPLOREDNS_WEBHOOK_TOKEN` | *(unset)* | Optional bearer token for webhook deliveries. When set, every webhook POST carries `Authorization: Bearer <token>`; pair it with the receiver's `RECEIVER_INGEST_TOKEN`. |
### Usage reporting
When `EXPLOREDNS_WEBHOOK_URL` is set, the server sends two JSON POSTs per
traversal, each with header `X-ExploreDNS-Event` naming the event:
- `start` — `{"event":"start","id","domain","query_type","all_roots","client_ip","started_at"}`
- `complete` — `{"event":"complete","id","domain","query_type","client_ip","started_at","done_at","duration_ms","status","error","result_count","summary"}`
where `summary` is the same grouped answers/statuses object returned by
`GET /api/traverse/{id}` and `error` is present only for failed jobs.
`client_ip` is the requester's IP (`Fly-Client-IP`, else the first
`X-Forwarded-For` entry, else the connection address). Delivery is
fire-and-forget: a 5-second timeout, one retry after 2 seconds, and failures
are logged without ever affecting the traversal or the API response. On
Fly.io, configure it (and the optional bearer token) as secrets rather than
in `fly.toml`:
```sh
fly secrets set EXPLOREDNS_WEBHOOK_URL=https://example.com/hook \
EXPLOREDNS_WEBHOOK_TOKEN=some-long-random-string
```
This repo ships a matching receiver for these events — see
[Usage telemetry receiver](#usage-telemetry-receiver).
---
## Deploying to Fly.io
The repo ships a [fly.toml](fly.toml) that builds `Dockerfile.web` and runs
the web server with scale-to-zero machines in `syd` (edit `app` /
`primary_region` to taste).
### First-time setup
```sh
flyctl auth login
flyctl apps create exploredns # match the app name in fly.toml
make deploy # flyctl deploy --remote-only
```
`make deploy-status` shows machine and health-check state. The app serves
the SPA at `https://<app>.fly.dev/` with `/api/health` as the health check.
Machine placement is imperative rather than part of `fly.toml`; the current
production topology is four regions — Sydney, Virginia, Singapore, and
London — so anycast wake-up behaviour can be observed from anywhere:
```sh
flyctl scale count 4 --region syd,iad,sin,lhr
```
`flyctl deploy` preserves existing machines and regions on redeploys.
`GET /api/health` reports which region served the request (`region` field,
present only on Fly), making the routing easy to observe:
```sh
curl -s https://exploredns.hansenits.com/api/health | jq -r .region
```
### Continuous deployment
`.gitea/workflows/deploy.yml` deploys on any `v*` tag push (or manual
dispatch). It needs a `FLY_API_TOKEN` repository secret:
```sh
flyctl tokens create deploy -x 999999h
```
### Releases
Pushing a `v*` tag triggers the full release pipeline:
1. `.gitea/workflows/release.yml` (`binaries` job) cross-compiles the CLI,
server, and receiver for linux/amd64, linux/arm64, darwin/amd64,
darwin/arm64 and windows/amd64, packages them as
`exploredns_<tag>_<os>_<arch>.tar.gz` (`.zip` on Windows) plus a
`SHA256SUMS` file, and attaches everything to the Gitea release for the
tag. Create the release with notes by hand before (or after) pushing
the tag — the workflow attaches assets to an existing release, creates
a bare one only when none exists, and skips already-attached assets so
re-runs are safe.
2. `.gitea/workflows/release.yml` (`docker` job) pushes
`gitea.hansenits.com.au/hits/exploredns-cli`, `…/exploredns-web`, and
`…/exploredns-receiver` images tagged `<tag>` and `latest`.
3. `.gitea/workflows/deploy.yml` deploys the web server to Fly.io.
All binaries are stamped with the tag via
`-ldflags "-X main.version=<tag>"`; check with `exploredns --version` or
`GET /api/health`. Local `make build` stamps from
`git describe --tags --always`.
### Notes
- Traversal traffic is outbound UDP/TCP port 53, which Fly machines allow;
upstream root discovery uses Fly's internal resolver via `/etc/resolv.conf`
and falls back to the built-in IANA root hints.
- The job timeout, job cap, per-IP rate limit, and same-origin CORS defaults
above are what make unauthenticated public exposure reasonable; tighten
`EXPLOREDNS_MAX_JOBS` or `EXPLOREDNS_RATE_LIMIT` if the app attracts
traffic.
---
### Text (default)
dnstraverse-style output: a header block (settings, initial root, query;
suppressed by `--quiet`), progress lines (`<refid> <server> (<ips>)` with
` -- resolving` and ` -- completed earlier (<refid>)` markers), a `Results:`
section of aggregated outcomes with probabilities (`Answer from`, `No glue
at`, `Lame referral from`, error/exception wording), and a `Summary Results:`
section grouping outcomes by status and answer content. `--show-servers`
adds the sorted list of servers encountered. Colour is used only when stdout
is a terminal and the `NO_COLOR` environment variable is unset.
### JSON (`--json`)
A single JSON document — `{domain, qtype, root, results, summary, servers}` —
emitted once at the end of the run with deterministic ordering. Each
aggregated outcome appears exactly once in `results`; `summary` groups
probabilities by status and by distinct answer RRset; `servers` is present
with `--show-servers`. Suitable for piping into `jq`.
---
## Usage telemetry receiver
`cmd/exploredns-receiver` is a small companion service that receives the
usage webhooks described above (`start`/`complete` events from
`EXPLOREDNS_WEBHOOK_URL`), stores them in MySQL or SQLite, and serves a
basic-auth-protected admin dashboard (`/admin`) plus JSON API
(`/admin/api/traversals`, `/admin/api/stats`) over the collected data. It
is a separate binary intended to run wherever you keep long-lived storage
(e.g. a home Kubernetes cluster) while the public web server stays
stateless.
Endpoints: `POST /webhook` (ingest, bearer-token protected when configured),
`GET /healthz` (liveness/readiness), `GET /admin` and `GET /admin/api/*`
(basic auth, always required).
### Configuration
| Variable | Default | Meaning |
|---|---|---|
| `RECEIVER_ADDR` | `:8080` | Listen address. |
| `RECEIVER_MYSQL_DSN` | *(unset)* | [go-sql-driver DSN](https://github.com/go-sql-driver/mysql#dsn-data-source-name) (`user:pass@tcp(host:3306)/dbname`). When set, events are stored in MySQL and the SQLite settings are ignored. |
| `RECEIVER_SQLITE_PATH` | `data/exploredns-receiver.db` | SQLite database path, used when no MySQL DSN is set (the container image defaults it to `/data/exploredns-receiver.db`). Parent directories are created automatically. |
| `RECEIVER_INGEST_TOKEN` | *(unset)* | When set, `POST /webhook` requires `Authorization: Bearer <token>`. Leave unset only on trusted networks. |
| `RECEIVER_ADMIN_USER` | `admin` | Basic-auth username for `/admin`. |
| `RECEIVER_ADMIN_PASSWORD` | *(required)* | Basic-auth password for `/admin`; the receiver refuses to start without it. |
Both storage backends share one portable schema; pure-Go drivers
(`modernc.org/sqlite`, `github.com/go-sql-driver/mysql`) keep the binary
CGO-free. SQLite is the zero-setup default; point `RECEIVER_MYSQL_DSN` at
an external MySQL when you want the data outside the pod/VM.
### Running with Docker
```sh
docker run -d --name exploredns-receiver \
-p 8080:8080 \
-v exploredns-receiver-data:/data \
-e RECEIVER_ADMIN_PASSWORD=change-me \
-e RECEIVER_INGEST_TOKEN=some-long-random-string \
gitea.hansenits.com.au/hits/exploredns-receiver:latest
```
The image stores SQLite data under the `/data` volume; add
`-e RECEIVER_MYSQL_DSN=...` to use MySQL instead.
### Running on Kubernetes
[deploy/k8s/receiver/](deploy/k8s/receiver/) contains commented template
manifests: a single-replica deployment (SQLite on a 1Gi PVC mounted at
`/data`, probes on `/healthz`), ClusterIP service, ingress with TLS
placeholders, and a secret template for the `RECEIVER_*` variables. Edit
the placeholder host/credentials, then:
```sh
kubectl apply -f deploy/k8s/receiver/
```
Keep one replica while on SQLite; MySQL removes that constraint.
### Pairing with the web server
Set the same token on both ends so the receiver only accepts events from
your server — e.g. on Fly.io:
```sh
fly secrets set EXPLOREDNS_WEBHOOK_URL=https://receiver.example.com/webhook \
EXPLOREDNS_WEBHOOK_TOKEN=some-long-random-string
# receiver side: RECEIVER_INGEST_TOKEN=some-long-random-string
```
---
## Comparison with dnstraverse
| Feature | dnstraverse (Ruby) | ExploreDNS (Go) |
|---|---|---|
| Language | Ruby | Go |
| Self-contained binary | No | Yes |
| All root servers | Yes | Yes |
| No-glue resolution | Yes | Yes |
| JSON output | No | Yes |
| Server fingerprinting | Yes | Yes |
| CNAME loop detection | Partial | Yes |
| Active development | Dormant | Active |
---
## Project Structure
```
cmd/exploredns/ CLI entry point and flag parsing
cmd/server/ HTTP API server entry point
cmd/exploredns-receiver/ Usage telemetry receiver entry point
internal/config/ Configuration types, validation, and usage text
internal/dns/ DNS query layer, root discovery, transport
internal/traverse/ Core traversal engine, referral resolution, caching
internal/fingerprint/ DNS server version fingerprinting (version.bind CHAOS)
internal/output/ Result formatting — text tree and JSON renderers
internal/integration/ End-to-end integration tests
internal/receiver/ Telemetry receiver: HTTP server, admin UI, event store
web/api/ HTTP handler, job store, SSE streaming, static assets
deploy/k8s/receiver/ Kubernetes manifest templates for the receiver
```
---
## Development
```sh
make build # compile CLI binary to bin/exploredns
make build-server # compile server binary to bin/exploredns-server
make build-receiver # compile telemetry receiver to bin/exploredns-receiver
make build-all # compile all three binaries
make test # run all unit and integration tests
make lint # run go vet
make clean # remove build artefacts
```
Run a single package's tests:
```sh
go test ./internal/traverse/...
```
---
## Architecture Overview
```
main → config.Parse → traverse.NewTraverser → traverse.Traverse
│
┌─────────▼──────────┐
│ Stack (BFS/DFS) │
│ Referral queue │
└─────────┬──────────┘
│ per referral
┌─────────▼──────────┐
│ processReferral │
│ ├─ resolveGlue │ (no-glue NS resolution)
│ └─ queryServer │ (dns.Query)
└─────────┬──────────┘
│
┌─────────────▼──────────────┐
│ Response classifier │
│ Answer / Referral / │
│ CNAME / NXDOMAIN / │
│ SERVFAIL / Error │
└─────────────┬──────────────┘
│
┌─────────▼──────────┐
│ output.Formatter │
│ (text | JSON) │
└────────────────────┘
```
The traversal engine (`internal/traverse`) maintains a work stack of
`Referral` objects. Each referral represents a single query to a single set
of nameservers. When a response contains further referrals, child `Referral`
objects are pushed onto the stack and processed in turn.
The `InfoCache` is used to store glue records discovered during traversal. In
fast mode (default), a single root cache is shared across all branches so that
glue discovered early is reused. In non-fast mode, each branch gets its own
independent cache.
---
## License
MIT — see [LICENSE](LICENSE).