Troubleshooting
View as MarkdownCommon errors and fixes when deploying or operating the Ory stack.
Where to look first
For most failures, the right place to start is the Ory pod logs:
kubectl logs -n ory deploy/kratos -f
kubectl logs -n ory deploy/hydra -f
kubectl logs -n ory deploy/ory-selfservice-ui -f
kubectl logs -n ory deploy/polis -f # when enable_polis = true
Hydra Maester (responsible for managing OAuth2Client CRDs):
kubectl logs -n ory deploy/hydra-hydra-maester -f
For login-time issues, tail Kratos, Hydra, and Polis simultaneously so you can see which component rejected the flow.
Symptom table
| Symptom | Likely cause | Fix |
|---|---|---|
Pods stuck in ImagePullBackOff |
License key JWT missing the ory entitlement, or expired |
Update license_key in tfvars, re-apply, restart the Ory pods |
curl https://polis.example.com times out, LB has zero endpoints |
Service selector doesn’t match polis pod labels | Confirm the LoadBalancer Service selector targets app.kubernetes.io/name=polis, instance=polis |
| TLS handshake fails on polis hostname | Polis serves plain HTTP internally; the Polis chart’s TLS sidecar terminates HTTPS | Check the Polis pod’s TLS sidecar container is running and the polis-tls certificate secret exists |
Polis SCIM endpoint URLs return http://localhost:5225/... |
EXTERNAL_URL env var not set on Polis |
Module sets it automatically from external_url; re-run terraform apply |
Polis logs OAuth server not configured correctly for openid flow, check if JWT signing keys are loaded |
OPENID_RSA_PRIVATE_KEY and OPENID_RSA_PUBLIC_KEY missing |
Module auto-generates and injects them; re-run terraform apply |
Polis logs "pkcs8" must be PKCS#8 formatted string |
RSA private key was PKCS#1 | Module uses the PKCS#8 form; re-run terraform apply |
Terraform apply fails on the Polis Helm release with failed to create patch: The order in patch list ... doesn't match $setElementOrder list |
The live Polis Deployment has a duplicated environment variable (older module versions set OPENID_REDIRECT_EXACT_MATCH twice), which Kubernetes can’t patch once the list changes |
Delete the Deployment (kubectl -n ory delete deployment polis) and re-run terraform apply, which recreates it; Polis’s data lives in its database, so nothing is lost |
First login fails with no matching authentication claim found in the JWT |
Hydra issued a token without identity claims. Common causes: the OAuth2 client has skipConsent: true, or a kratos_helm_values override dropped the module’s oidc/saml registration session hooks |
Keep skipConsent: false on the Materialize client (the module default). The consent handler injects the email and groups claims. If you override kratos_helm_values, keep the registration after hooks for oidc and saml |
"Couldn't fetch XML data" when registering a Polis SAML connection |
The IdP’s metadata URL is gated by API auth | Post rawMetadata=<XML> to Polis instead of metadataUrl=... |
| “Sign in via SAML” button missing on Kratos login | Cached login flow from before the polis provider was added | Hard refresh or open a new incognito session |
| Materialize Console reaches the login screen but balancerd times out | DNS or cert SAN mismatch | Confirm the balancerd hostname A record resolves and is in the cert SAN list (balancerd_extra_dns_names) |
| User logs in but can’t run any SQL | JIT role created with no privileges | Run GRANT <role> TO "user@email" as mz_system |
| SCIM “Test Connector Configuration” passes but no users push | Existing assignments don’t backfill when SCIM is enabled after-the-fact | In Okta, unassign + reassign the user, or push profile updates from the people side |
terraform destroy hangs on kubernetes_namespace.ory |
OAuth2Client finalizer not cleared because Hydra Maester is torn down before processing it | kubectl patch oauth2client materialize-oauth2-client -n ory --type=json -p='[{"op":"remove","path":"/metadata/finalizers"}]' then re-run destroy |
Detailed walkthroughs
Hydra returns invalid_client on the token endpoint
Usually means the OAuth2Client CRD didn’t reconcile against Hydra (so
the client_id Materialize is using doesn’t exist in Hydra’s
database).
Check the CRD:
kubectl get oauth2client -n ory materialize-oauth2-client -o yaml
Look for the status.reconciliationError field. Common errors:
- “Hydra admin unreachable” – Hydra Maester can’t connect to Hydra’s
admin port. Confirm
hydra-admin.ory.svc.cluster.local:4445resolves from inside the cluster. - “duplicate client name” – a stale OAuth2Client from a previous apply exists in Hydra’s DB. Delete the CRD, wait for Hydra Maester to drop the Hydra-side record, then re-apply.
Check Hydra Maester logs:
kubectl logs -n ory deploy/hydra-hydra-maester --tail=100
Kratos selfservice UI shows a blank login page
Almost always a TLS or DNS issue between the browser and Kratos. Open your browser’s network tab and inspect the requests:
- 502/504 on
/self-service/login/browser→ Kratos isn’t reachable from the UI pod. Check Kratos pod status. - TLS error on the redirect target → cert not provisioned yet, or the
hostname DNS record isn’t propagated. Run
kubectl get certificate -Aand confirmkratos-tlsisREADY=True. CORSerror → the Hydracors_allowed_originsdoesn’t include the console hostname. Check thehydraHelm release values.
Polis SAML flow returns server_error: <random-name>
Polis encodes errors with a random nickname for the log entry (e.g.
curve_tourist_bean). The real error is in the Polis pod logs:
kubectl logs -n ory deploy/polis --tail=200 | grep -B2 -A5 "error\|Error"
Common causes:
"pkcs8" must be PKCS#8 formatted string→ see symptom table- IdP metadata doesn’t include the SAML signing cert → re-export the
metadata XML from the IdP and update the Polis connection with
rawMetadata - Audience mismatch → confirm the IdP’s SAML app audience is set to
https://saml.boxyhq.com
Console login loops back to the IdP indefinitely
Usually a redirect URI mismatch. The redirect URI registered with your IdP must exactly match what Kratos sends:
https://<your-kratos-hostname>/self-service/methods/oidc/callback/<id>
where <id> is the entry in upstream_identity_providers. Trailing slashes
matter. Update the redirect URI in the IdP to match.
If the redirect URI is correct, check the oidc_audience system
parameter on the Materialize side. It must include the OAuth2 client_id
Hydra Maester generated:
kubectl get secret -n ory materialize-oauth2-client \
-o jsonpath='{.data.CLIENT_ID}' | base64 -d
Compare with the Materialize CR’s system_parameters.oidc_audience.
Inspecting the JWT
To see exactly which claims a Hydra-issued token carries, sign in through the
browser, then in DevTools grab the id_token (Application → Cookies, or the
OAuth callback response in the Network tab) and decode the middle segment:
echo '<paste-JWT>' | cut -d. -f2 | base64 -d 2>/dev/null | jq
Look for email, iss (should match your ory_hydra_fqdn), and any custom
claims you configured (groups, etc.).
JWT is missing a custom claim (e.g. groups) even though the IdP is sending it
Kratos’s OIDC jsonnet mapper exposes standard OpenID claims (email, sub,
aud, iss, preferred_username, etc.) as top-level keys on the claims
object, but any non-standard claim (groups, department, tenant_id, …)
lives under claims.raw_claims.
If the mapper reads claims.groups, it will silently return null or an empty
default even when the IdP token clearly contains groups. Read from
claims.raw_claims (with a fallback for providers that flatten):
local claims = std.extVar('claims');
local raw = if std.objectHas(claims, 'raw_claims') then claims.raw_claims else {};
local groups_from(src) = if std.objectHas(src, 'groups') then src.groups else [];
{
identity: {
traits: {
email: claims.email,
groups: if std.length(groups_from(claims)) > 0 then groups_from(claims) else groups_from(raw),
},
},
}
This is the pattern the module ships. When adding new IdP-specific custom
claims to the mapper, always check raw_claims first.
Okta reports “Invalid Base URL for the SCIM Connector”
The SCIM connector base URL Polis returns
(<polis-hostname>/api/scim/v2.0/<directoryId>) must be entered into Okta
with a trailing slash. Without it, Okta’s client-side validation rejects
the URL before making any HTTP request. Add / at the end and re-test.
Okta’s “Test Connector Configuration” fails with “Error authenticating: null”
Okta calls the SCIM base URL from its own cloud, so the host in that URL must
be reachable from the internet (restricted to your IdP’s egress ranges if you
like). If Polis sits behind an internal load balancer or a private network, the
test fails with this unhelpful message even though the token is correct. Use a
hostname that resolves to a publicly reachable Polis endpoint for the SCIM base
URL, keeping the same /api/scim/v2.0/<directoryId>/ path and token.
Users don’t appear in Polis’s directory after assigning a group
Assigning a group to the SAML app under Assignments normally triggers user provisioning for each member, but Okta silently skips users whose profile is missing an attribute the SCIM mapping requires (email in particular).
Diagnose and unstick:
-
Check the app’s Assignments tab. Each assigned user has a push status column. A red icon means provisioning failed; hover for the reason. If it says the user “was assigned this application before Provisioning was enabled”, click Provision User at the top of the tab to provision all pending users.
-
Force a push manually: Assignments tab → Assign → Assign to People and add the user by email.
-
Verify in Polis:
curl -s -H "Authorization: Api-Key $POLIS_API_KEY" \ "https://<your-polis-hostname>/api/v1/dsync/users?directoryId=$DIRECTORY_ID" | jq .
Note: pushing a group via the Push Groups tab creates the group entity in Polis but does not push its members as SCIM users. Members are provisioned via Assignments.
A specific cloud is hanging during destroy
See the cloud-specific notes:
- Azure: usually the OAuth2Client finalizer (see the symptom table above)
- AWS: the AWS Load Balancer Controller can race with namespace deletion. See AWS-specific notes
- GCP: the GKE master IP allocator can be slow; usually patience is the fix
When to escalate
If the failure doesn’t match any of the above and the Ory pod logs don’t surface a clear cause, file an issue with:
- The pod logs (
kubectl logs -n ory deploy/<component>) from the affected component plus the two it talks to - The Hydra OAuth2Client CRD YAML (
kubectl get oauth2client -n ory materialize-oauth2-client -o yaml) - The output of
kubectl get all,certificate,oauth2client -n ory - The cloud (Azure / GCP / AWS), Materialize version, and Ory chart version pinned in your tfvars