Docs

Pro failover

Pro gives you a dedicated address in France (GRA) and one in Germany (LIM). Your jobs try GRA first and use LIM when GRA does not answer. The switch happens in your job, on your side. We do not switch traffic for you.

What this is, and is not

Client-side failover: a job chooses an endpoint when it starts. There is no automatic failover on our side, and the pilot is best effort with no SLA. A connection that is open when an endpoint stops working breaks, and your job has to retry.

  1. How Pro is set up
  2. When to switch
  3. In GitHub Actions
  4. In a shell script
  5. With WireGuard
  6. Practice

How Pro is set up

  • Two dedicated IPv4 addresses: one in Gravelines, France (GRA), one in Limburg, Germany (LIM).
  • Each address is its own endpoint. It has its own proxy login, its own HTTPS proxy port and its own WireGuard configuration. A login for one does not work on the other.
  • Your allowlist applies to both addresses. See Allowlist.
  • The destination has to allowlist both addresses. If it lists only one, a switch turns an outage on our side into a refusal on theirs.

We suggest GRA as the primary and LIM as the standby, but the order is yours. Put first the one that is closer to the destinations you call.

When to switch

Switch when the endpoint itself is not usable, and only then.

What you seeSwitch?Why
Connection to the proxy times out or is refusedYesThe endpoint or the path to it is down
WireGuard shows no handshakeYesThe tunnel does not come up
Proxy answers 503Yes, after a short retryConnection limit reached, or the proxy is restarting
Proxy answers 407NoThe login is wrong or was rotated. Fix the credentials; the other endpoint has a different login
Proxy answers 403NoThe destination is not on your allowlist. The other address is refused the same way
The destination refuses youNoCheck that it allowlists the address you used, then talk to its owner

In GitHub Actions

The GitHub Action takes both endpoints. It probes the primary with your login, uses it when it works, and otherwise tries the standby with the standby's login. It warns when it took the standby and reports the choice in the endpoint output. It is more forgiving than the table above: it moves on to the standby after any failure of the primary, a rejected login included (each endpoint has its own login), and when no endpoint works it fails the job and lists why each one failed. Treat the warning as something to fix, not as normal operation. With verify-url and expected-ip it also checks that the address the destination sees is the one you expect for that endpoint.

.github/workflows/sync.yml
      - name: Route the job through the dedicated IP
        id: egress
        uses: wegvon/penduses-egress@v1
        with:
          mode: proxy
          # Primary (GRA) first, standby (LIM) after the comma
          endpoints: ${{ vars.PENDUSES_ENDPOINTS }}
          username: ${{ secrets.PENDUSES_GRA_USER }}
          password: ${{ secrets.PENDUSES_GRA_PASSWORD }}
          standby-username: ${{ secrets.PENDUSES_LIM_USER }}
          standby-password: ${{ secrets.PENDUSES_LIM_PASSWORD }}
          verify-url: https://ip-echo.partner.example/ip
          expected-ip: 203.0.113.25,192.0.2.25

      - name: Call the partner API from the dedicated IP
        run: |
          . "${{ steps.egress.outputs.env-file }}"
          curl -sS --fail https://api.partner.example/v1/sync

In a shell script

Ask each endpoint, in order, to open a tunnel to a service on your allowlist, and export the proxy of the first one that agrees. The script reads the logins from the environment, keeps them off the command line, and stops when no endpoint works. It switches only when the endpoint does not answer or is overloaded. A 407 or 403 is your login or your allowlist, which the other endpoint would not fix, so the script stops and says so.

choose-egress.sh
#!/usr/bin/env bash
set -euo pipefail

# HOST HTTP_PORT USER PASSWORD, primary first. HOST is the node's DNS name: the TLS certificate is issued for it. The logins come from your secret manager.
ENDPOINTS=(
  "198-51-100-10.sslip.io 20002 $GRA_USER $GRA_PASSWORD"
  "198-51-100-20.sslip.io 20012 $LIM_USER $LIM_PASSWORD"
)
CHECK_URL=https://ip-echo.partner.example/ip

# Print the status the proxy gave to CONNECT for CHECK_URL (000: no answer).
# The endpoint comes from the variables host, port, user and password.
connect_status() {
  curl --config - --silent --output /dev/null --write-out '%{http_connect}' \
      --connect-timeout 5 --max-time 15 "$CHECK_URL" <<EOF || true
proxy = "https://$host:$port"
proxy-user = "$user:$password"
proxytunnel
EOF
}

chosen=""
for endpoint in "${ENDPOINTS[@]}"; do
  read -r host port user password <<<"$endpoint"
  status=$(connect_status)
  case "$status" in
    200)
      export HTTPS_PROXY="https://$user:$password@$host:$port"
      chosen="$host"
      break
      ;;
    403|407)
      echo "endpoint $host answered $status: check the login and the allowlist" >&2
      exit 1
      ;;
    *)
      echo "endpoint $host is not usable (status $status), trying the next one" >&2
      ;;
  esac
done

if [[ -z "$chosen" ]]; then
  echo "no egress endpoint is usable" >&2
  exit 1
fi
echo "using endpoint $chosen"

Source it (. ./choose-egress.sh) in the job that needs the dedicated address. The check URL must be on your allowlist.

With WireGuard

Keep one configuration per location, named after it, and bring up one at a time. Both routes cover the same destinations, so the two tunnels must not be up together.

choose-tunnel.sh
#!/usr/bin/env bash
set -euo pipefail

# Wait up to 10 seconds for the first handshake on the interface in $iface
handshake() {
  for _ in 1 2 3 4 5 6 7 8 9 10; do
    # latest-handshakes lists "peer  epoch"; 0 means no handshake yet
    if sudo wg show "$iface" latest-handshakes | cut -f2 | grep -qv '^0$'; then
      return 0
    fi
    sleep 1
  done
  return 1
}

active=""
for iface in gra lim; do
  if sudo wg-quick up "./$iface.conf" && handshake; then
    active="$iface"
    break
  fi
  sudo wg-quick down "./$iface.conf" || true
done

if [[ -z "$active" ]]; then
  echo "no tunnel came up" >&2
  exit 1
fi
echo "tunnel $active is up"

The GitHub Action does the same in WireGuard mode when you give it standby-wireguard-config, and takes the tunnel down when the job ends.

Practice

  • Test the standby before you need it. Run a scheduled job that uses the standby endpoint on purpose and checks the source address. An untested standby fails at the worst moment, for example because the destination never allowlisted it.
  • Make jobs retry. The choice is made when the step starts. A long job should repeat the choice, or retry its own calls, when a connection breaks.
  • Expect different behaviour. The two locations are in two countries. Latency differs, and a destination that decides by country or geolocation database may treat them differently.
  • Keep the logins apart. Store each endpoint's login under its own secret name, so a rotation of one does not touch the other.
  • Watch which endpoint you use. Log the choice. A job that quietly runs on the standby for weeks is a problem waiting to be noticed.

Pro is best effort with no SLA. To report an endpoint that is down, email support@penduses.com with your workspace name and the time (UTC). Do not send passwords or keys.