# Kong 1.4, K8S, DB-less, 504 Gateway timeout

**URL:** <https://discuss.konghq.com/t/kong-1-4-k8s-db-less-504-gateway-timeout/4956>\
**Category:** General\
**Tags:** kubernetes, kong-gateway\
**Created:** [November 20, 2019, 3:37pm UTC](https://discuss.konghq.com/t/kong-1-4-k8s-db-less-504-gateway-timeout/4956 "2019-11-20T15:37:56Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![abenitovsc](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.konghq.com/abenitovsc/32/1521_2.png) [@abenitovsc](https://discuss.konghq.com/u/abenitovsc)\
**Post date:** [November 20, 2019, 3:37pm UTC](https://discuss.konghq.com/t/kong-1-4-k8s-db-less-504-gateway-timeout/4956/1 "2019-11-20T15:37:56Z")

</div>

Hi all,

We have 3 Kong replicas deployed in Kubernetes from stable/kong 0.19.1 Helm chart). We have upgraded Kong to 1.4 using the alpine images. Kong is configured in DB-less mode.

We have checked that when the edge load balancer points to a particular replica, for a particular route, the service is not responding, returning 504 Gateway timeout after 60 secs. 2/3 response KO and 1/3 OK.

When accesing through a different replica there is no issue.

Next picture shows how requests balanced to pod “stingray-kong-6bf98f84d6-q7vrn” response properly with a 200 and requests to the other two pods fail with the client request closed (499)

 ![39](https://canada1.discourse-cdn.com/flex036/uploads/konghq/original/2X/5/58612dd87f8769b606b9a972915c627b78444625.png)

these behaviour cannot be reproduced requesting other apps and three nodes serve request correctly.

On the other hand we can see that the pod which serves properly that route has different memory consumption.

 ![19](https://canada1.discourse-cdn.com/flex036/uploads/konghq/original/2X/e/e6cd5db4c958d42df0ffa6c8c02cf5f99e21e65c.png)

the default Kong ingress configuration is used and no plugins for that Ingress.

Updated with more info:

- Restarting Kong deployments the issue dissapears for some hours.
- if the upstream is removed (deleting the pod) the issue persists. Restarting the upstream seems not working. It seems there is no a issue in the upstream, Kong might no be able to proxy the request.
- Deleting the ingress and creating it again doesn’t fix the issue.

---

<div class="post-metadata">

**Author:** ![hbagdi](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.konghq.com/hbagdi/32/562_2.png) [@hbagdi](https://discuss.konghq.com/u/hbagdi)\
**Post date:** [November 21, 2019, 7:26pm UTC](https://discuss.konghq.com/t/kong-1-4-k8s-db-less-504-gateway-timeout/4956/2 "2019-11-21T19:26:12Z")

</div>

Interesting.

Do you see any error in the controller of the two pods that were not able to serve requests?

---

<div class="post-metadata">

**Author:** ![abenitovsc](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.konghq.com/abenitovsc/32/1521_2.png) [@abenitovsc](https://discuss.konghq.com/u/abenitovsc)\
**Post date:** [November 21, 2019, 7:51pm UTC](https://discuss.konghq.com/t/kong-1-4-k8s-db-less-504-gateway-timeout/4956/3 "2019-11-21T19:51:52Z")

</div>

Nothing!

Currently one of them has been deleted and Kong retrieves 503 (Right behavior) and 504 alternatively.  
We can access directly with port-forwarding without any issue.

Attached a recent evidence made with Apachebench. Kong pods “-7rw…” and “-2cv…” retrieve 504, “-dnx…” retrieves 503.

 ![41](https://canada1.discourse-cdn.com/flex036/uploads/konghq/original/2X/3/3ec3d5225fac957769970bbd471f285a247a8521.png)

It seems that unstable pods are not releasing memory after 17:30

 ![29](https://canada1.discourse-cdn.com/flex036/uploads/konghq/original/2X/2/240ea792671e81f3e07f6306733df4128cf28843.png)

---

<div class="post-metadata">

**Author:** ![hbagdi](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.konghq.com/hbagdi/32/562_2.png) [@hbagdi](https://discuss.konghq.com/u/hbagdi)\
**Post date:** [November 21, 2019, 8:19pm UTC](https://discuss.konghq.com/t/kong-1-4-k8s-db-less-504-gateway-timeout/4956/4 "2019-11-21T20:19:03Z")

</div>

What is your k8s environment and version?  
The issue doesn’t seem to be with Kong itself but with Kubernetes or the load-balancer in-front of Kong.

Kong is responding with 499, meaning the load balancer closed the connection before Kong could send back the response.

Some things to ensure:

1. Is this issue happening with specific k8s worker nodes? Kong might not be able to reach the services in the rest of the cluster due a networking problem.
2. The Load-Balancer is able to reach Kong pods in other zones if you are running in a cloud provider.
3. All the worker nodes have proper networking configured and are not out of sync in any ways.

---

<div class="post-metadata">

**Author:** ![abenitovsc](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.konghq.com/abenitovsc/32/1521_2.png) [@abenitovsc](https://discuss.konghq.com/u/abenitovsc)\
**Post date:** [November 21, 2019, 8:35pm UTC](https://discuss.konghq.com/t/kong-1-4-k8s-db-less-504-gateway-timeout/4956/5 "2019-11-21T20:35:39Z")

</div>

Thank you,

- Each replica (3) is deployed in a different node.
- Currently there is no target upstream, it is not deployed.
- Load balancer discarded, I have just tried requesting directly to each pod (with port forwarding) using curl and only “-dnx” is responding properly.
- Kubernetes version is 1.11
- We have a lot of environments deployed in the cluster and these are the only two pods with the issue.
- One pod is Jenkins and the other is this: [https://hub.docker.com/r/brndnmtthws/nginx-echo-headers/](https://hub.docker.com/r/brndnmtthws/nginx-echo-headers/)

---

<div class="post-metadata">

**Author:** ![hbagdi](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.konghq.com/hbagdi/32/562_2.png) [@hbagdi](https://discuss.konghq.com/u/hbagdi)\
**Post date:** [November 21, 2019, 9:34pm UTC](https://discuss.konghq.com/t/kong-1-4-k8s-db-less-504-gateway-timeout/4956/6 "2019-11-21T21:34:29Z")

</div>

1. Do you mean that only two services that are being proxies by Kong have this issue or is it two Kong pods that have this issue?

2. After LB is removed, what error do you get back from Kong when you curl it directly over the port-forwarding tunnel?

---

<div class="post-metadata">

**Author:** ![abenitovsc](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.konghq.com/abenitovsc/32/1521_2.png) [@abenitovsc](https://discuss.konghq.com/u/abenitovsc)\
**Post date:** [November 21, 2019, 10:07pm UTC](https://discuss.konghq.com/t/kong-1-4-k8s-db-less-504-gateway-timeout/4956/7 "2019-11-21T22:07:12Z")

</div>

Yes, only these two services have the issue, the others work perfectly.

I have repeated the test asking resources to Jenkins directly to a Kong pod (without ELB) and this is the result:

```
> 
> 2019/11/21 21:42:46 [error] 36#0: *182367757 [lua] init.lua:800: balancer(): failed to retry the dns/balancer resolver for jenkins.ci.svc' with: dns server error: 100 cache only lookup failed while connecting to upstream, client: 127.0.0.1, server: kong, request: "GET /testIssue3 HTTP/1.1", upstream: "http://100.67.34.45:80/testIssue3", host: " ****" 
> 
> 
> 2019/11/21 21:42:46 [error] 36#0: *182367757 [lua] balancer.lua:900: balancer_execute(): DNS resolution failed: dns server error: 100 cache only lookup failed. Tried: ["(short)jenkins.ci.svc:(na) - cache-miss","jenkins.ci.svc.ingress.svc.cluster.local:33 - cache only lookup failed/dns server error: 100 cache only lookup failed","jenkins.ci.svc.svc.cluster.local:33 - cache only lookup failed/dns server error: 100 cache only lookup failed","jenkins.ci.svc.cluster.local:33 - cache only lookup failed/dns server error: 100 cache only lookup failed","jenkins.ci.svc.eu-west-1.compute.internal:33 - cache only lookup failed/dns server error: 100 cache only lookup failed","jenkins.ci.svc:33 - cache only lookup failed/dns server error: 100 cache only lookup failed","jenkins.ci.svc.ingress.svc.cluster.local:1 - cache only lookup failed/dns server error: 100 cache only lookup failed","jenkins.ci.svc.svc.cluster.local:1 - cache only lookup failed/dns server error: 100 cache only lookup failed","jenkins.ci.svc.cluster.local:1 - cache only lookup failed/dns server error: 100 cache only lookup failed","jenkins.ci.svc.eu-west-1.compute.internal:1 - cache only lookup failed/dns server error: 100 cache only lookup failed","jenkins.ci.svc:1 - cache only lookup failed/dns server error: 100 cache only lookup failed","jenkins.ci.svc.ingress.svc.cluster.local:5 - cache only lookup failed/dns server error: 100 cache only lookup failed","jenkins.ci.svc.svc.cluster.local:5 - cache only lookup failed/dns server error: 100 cache only lookup failed","jenkins.ci.svc.cluster.local:5 - cache only lookup failed/dns server error: 100 cache only lookup failed","jenkins.ci.svc.eu-west-1.compute.internal:5 - cache only lookup failed/dns server error: 100 cache only lookup failed","jenkins.ci.svc:5 - cache only lookup failed/dns server error: 100 cache only lookup failed"] while connecting to upstream, client: 127.0.0.1, server: kong, request: "GET /testIssue3 HTTP/1.1", upstream: "http://100.67.34.45:80/testIssue3", host: " *****" t
> 
> 
> 2019/11/21 21:42:46 [error] 36#0: *182367757 upstream timed out (110: Operation timed out) while connecting to upstream, client: 127.0.0.1, server: kong, request: "GET /testIssue3 HTTP/1.1", upstream: "http://100.67.34.45:80/testIssue3"

```

I have tried a nslookup within the Kong pod and Jenkins.ci.svc is resolved correctly to 100.67.34.45.  
The Jenkins Kubernetes service has the following config:

> kind: Service  
> metadata:  
> annotations:  
> [kubectl.kubernetes.io/last-applied-configuration:](http://kubectl.kubernetes.io/last-applied-configuration:) |  
> {“apiVersion”:“v1”,“kind”:“Service”,“metadata”:{“annotations”:{},“labels”:{“name”:“jenkins”},“name”:“jenkins”,“namespace”:“ci”},“spec”:{“ports”:[{“name”:“jenkins-http”,“port”:8080,“protocol”:“TCP”,“targetPort”:8080}],“selector”:{“name”:“jenkins”}}}  
> creationTimestamp: “2019-07-30T12:56:37Z”  
> labels:  
> name: jenkins  
> name: jenkins  
> namespace: ci  
> resourceVersion: “50311438”  
> selfLink: /api/v1/namespaces/ci/services/jenkins  
> uid: 777266da-b2c9-11e9-9c10-027d1765881c  
> spec:  
> clusterIP: 100.67.34.45  
> ports:
> 
> - name: jenkins-http  
> port: 8080  
> protocol: TCP  
> targetPort: 8080  
> selector:  
> name: jenkins  
> sessionAffinity: None  
> type: ClusterIP  
> status:  
> loadBalancer: {}

why do the traces show the upstream with port 80?  
if I get the upstream from the Kong API it display the right Jenkins pod IP with right port.

```
|data||
|---|---|
|0||
|created_at|1574373014.286|
|upstream||
|id|"e9753ef9-c994-59e5-b348-4a132b793242"|
|id|"9f8052ed-8645-5129-8eab-84f168238129"|
|target|"100.96.8.221:8080"|
|weight|100|

NAME ENDPOINTS AGE
jenkins 100.96.8.221:8080 114d
jenkins-agent 100.96.8.221:50000 114d

```

---

<div class="post-metadata">

**Author:** ![hbagdi](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.konghq.com/hbagdi/32/562_2.png) [@hbagdi](https://discuss.konghq.com/u/hbagdi)\
**Post date:** [November 21, 2019, 10:29pm UTC](https://discuss.konghq.com/t/kong-1-4-k8s-db-less-504-gateway-timeout/4956/8 "2019-11-21T22:29:43Z")

</div>

Can you share the Ingress resource for Jenkins?

By any chance, do you have the [ingress.kubernetes.io/service-upstream](http://ingress.kubernetes.io/service-upstream) annotation applied anywhere?

---

<div class="post-metadata">

**Author:** ![abenitovsc](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.konghq.com/abenitovsc/32/1521_2.png) [@abenitovsc](https://discuss.konghq.com/u/abenitovsc)\
**Post date:** [November 22, 2019, 8:12am UTC](https://discuss.konghq.com/t/kong-1-4-k8s-db-less-504-gateway-timeout/4956/9 "2019-11-22T08:12:36Z")

</div>

Yes!

```
apiVersion: extensions/v1beta1
kind: Ingress
metadata:
  annotations:
    configuration.konghq.com: jenkins-ingress
    kubectl.kubernetes.io/last-applied-configuration: |
      {"apiVersion":"extensions/v1beta1","kind":"Ingress","metadata":{"annotations":{"kubernetes.io/ingress.class":"kong"},"name":"jenkins-ingress","namespace":"ci"},"spec":{"rules":[{"host":" ****","http":{"paths":[{"backend":{"serviceName":"jenkins","servicePort":"jenkins-http"}}]}}]}}
    kubernetes.io/ingress.class: kong
  creationTimestamp: "2019-11-15T14:08:19Z"
  generation: 3
  name: jenkins-ingress
  namespace: ci
  resourceVersion: "51274360"
  selfLink: /apis/extensions/v1beta1/namespaces/ci/ingresses/jenkins-ingress
  uid: 607ae6ed-07b1-11ea-9c10-027d1765881c
spec:
  rules:
  - host: *****
    http:
      paths:
      - backend:
          serviceName: jenkins
          servicePort: jenkins-http
status:
  loadBalancer:
    ingress:
    - hostname: ****

```

We are not using service-upstream annotations.

And this is the KongIngress resource content:

```
proxy:
  protocol: http
route:
  connect_timeout: 20000
  preserve_host: false
  regex_priority: 1
  strip_path: false

```

There is another interesting discussion here. “proxy” and “route” config works perfectly within the KongIngress CRD but “upstreams” doesn’t. We tried to update the http\_path and http\_status for the healtchecks but Kong admin API returns always the default upstreams configuration.

---

<div class="post-metadata">

**Author:** ![hbagdi](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.konghq.com/hbagdi/32/562_2.png) [@hbagdi](https://discuss.konghq.com/u/hbagdi)\
**Post date:** [November 22, 2019, 9:15pm UTC](https://discuss.konghq.com/t/kong-1-4-k8s-db-less-504-gateway-timeout/4956/10 "2019-11-22T21:15:08Z")

</div>

Okay. Nothing interesting here.  
Kong should not be doing a DNS lookup for jenkins.ci.svc hostname. That should be coming form Upstreams.

Do you see an upstream in Kong’s ADmin API associated with `jenkins.ci.svc` (and also targets)?

---

<div class="post-metadata">

**Author:** ![abenitovsc](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.konghq.com/abenitovsc/32/1521_2.png) [@abenitovsc](https://discuss.konghq.com/u/abenitovsc)\
**Post date:** [November 22, 2019, 10:39pm UTC](https://discuss.konghq.com/t/kong-1-4-k8s-db-less-504-gateway-timeout/4956/11 "2019-11-22T22:39:23Z")

</div>

Yes! upstream and targets are discovered correctly by Kong. However we get the same issue with and without targets.

First picture (Nginx traces) In this topic shows 200 responses and 499 when pod is running and target discovered.

Third picture, also Nginx traces, shows 503 response and 499 when pod is not running.

Currently with the pod running target matches the Pod IP and port.

```
{
  "created_at": 1574461462.702,
  "upstream": {
    "id": "e9753ef9-c994-59e5-b348-4a132b793242"
  },
  "id": "97a6c375-0209-516a-948e-964bacce09f8",
  "target": "100.96.8.25:8080",
  "weight": 100,
  "health": "HEALTHCHECKS_OFF"
}
```

---

<div class="post-metadata">

**Author:** ![abenitovsc](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.konghq.com/abenitovsc/32/1521_2.png) [@abenitovsc](https://discuss.konghq.com/u/abenitovsc)\
**Post date:** [November 25, 2019, 4:39pm UTC](https://discuss.konghq.com/t/kong-1-4-k8s-db-less-504-gateway-timeout/4956/12 "2019-11-25T16:39:39Z")

</div>

Hi!

I have realised that Grafana Kong dashboard shows the following result:

 ![06](https://canada1.discourse-cdn.com/flex036/uploads/konghq/original/2X/f/f751f79f74acfb696c69b30967926bdfd50024da.png)

What is happening with the kong-process-events?

---

<div class="post-metadata">

**Author:** ![hbagdi](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.konghq.com/hbagdi/32/562_2.png) [@hbagdi](https://discuss.konghq.com/u/hbagdi)\
**Post date:** [November 26, 2019, 12:07am UTC](https://discuss.konghq.com/t/kong-1-4-k8s-db-less-504-gateway-timeout/4956/13 "2019-11-26T00:07:40Z")

</div>

You can ignore the `kong_process_events` being full. It is allocated to the full amount but is not in complete use. At least, that’s the case in most DB-less deployments.

Regarding the original problem, I’ve no clue of why Kong is trying to resolve the DNS name `jenkins.ci.svc`.  
Can you make sure that the route -\> service -\> upstream -\> target connections hold up correctly?  
Meaning, the service \<-\> upstream connection is setup correctly?  
The way those two are connected are by the `host` property of the service object should be same as the `name` property of the upstream object.

---

<div class="post-metadata">

**Author:** ![abenitovsc](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.konghq.com/abenitovsc/32/1521_2.png) [@abenitovsc](https://discuss.konghq.com/u/abenitovsc)\
**Post date:** [November 27, 2019, 11:47am UTC](https://discuss.konghq.com/t/kong-1-4-k8s-db-less-504-gateway-timeout/4956/14 "2019-11-27T11:47:43Z")

</div>

Following the changes described in this topic, I have applied them and I will response with the feedback.

> [@Kong v 1.4.0 - Getting 504s when adding plugin to route via REST (SOLVED)](https://discuss.konghq.com/t/kong-v-1-4-0-getting-504s-when-adding-plugin-to-route-via-rest-solved/5000):
>
> 504 when adding plugin to route via REST Wondering if anyone else is hitting the issue too? This takes forever via curl and also used konga GUI to test the same thing, adding a plugin to an existing route is actually working, however the response is always a timeout, and so I have to add plugin, let it time out, move on to the next one. Kong is using DB and is running in docker. Same issue is not apparent in kong 1.3x eg curl -s -X POST [https://kong-admin/routes/](https://kong-admin/routes/){route id}/plugins -H ‘Cont…

---

<div class="post-metadata">

**Author:** ![abenitovsc](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.konghq.com/abenitovsc/32/1521_2.png) [@abenitovsc](https://discuss.konghq.com/u/abenitovsc)\
**Post date:** [November 28, 2019, 7:49am UTC](https://discuss.konghq.com/t/kong-1-4-k8s-db-less-504-gateway-timeout/4956/15 "2019-11-28T07:49:18Z")

</div>

These changes didn’t fix my issue, i will try to investigate more.

---

<div class="post-metadata">

**Author:** ![abenitovsc](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.konghq.com/abenitovsc/32/1521_2.png) [@abenitovsc](https://discuss.konghq.com/u/abenitovsc)\
**Post date:** [December 9, 2019, 2:00pm UTC](https://discuss.konghq.com/t/kong-1-4-k8s-db-less-504-gateway-timeout/4956/16 "2019-12-09T14:00:06Z")

</div>

After one week working with Kong 1.4.1, i confirm that new version solves this issue.

---

<div class="post-metadata">

**Author:** ![hbagdi](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.konghq.com/hbagdi/32/562_2.png) [@hbagdi](https://discuss.konghq.com/u/hbagdi)\
**Post date:** [December 9, 2019, 4:46pm UTC](https://discuss.konghq.com/t/kong-1-4-k8s-db-less-504-gateway-timeout/4956/17 "2019-12-09T16:46:01Z")

</div>

Glad to finally hear this!  
Thank you for leaving this message here.

---

<div class="post-metadata">

**Author:** ![HYUN\_SUK\_JUNG](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.konghq.com/hyun_suk_jung/32/1122_2.png) [@HYUN\_SUK\_JUNG](https://discuss.konghq.com/u/HYUN_SUK_JUNG)\
**Post date:** [January 22, 2020, 11:51am UTC](https://discuss.konghq.com/t/kong-1-4-k8s-db-less-504-gateway-timeout/4956/18 "2020-01-22T11:51:29Z")

</div>

Hello abenitovsc  
I have same issue. Is it clear the issue?  
Please let me know the version of kong

“balancer.lua:917: balancer\_execute(): DNS resolution failed: dns server error: 100 cache only lookup failed.”

I use kong 1.4.2 ingress 0.6.0 DB-less

---

<div class="post-metadata">

**Author:** ![abenitovsc](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.konghq.com/abenitovsc/32/1521_2.png) [@abenitovsc](https://discuss.konghq.com/u/abenitovsc)\
**Post date:** [January 22, 2020, 12:15pm UTC](https://discuss.konghq.com/t/kong-1-4-k8s-db-less-504-gateway-timeout/4956/19 "2020-01-22T12:15:41Z")

</div>

Hello,

the real issue was the 504 HTTP code, can you provide metrics for the Kong pods?

---

<div class="post-metadata">

**Author:** ![HYUN\_SUK\_JUNG](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.konghq.com/hyun_suk_jung/32/1122_2.png) [@HYUN\_SUK\_JUNG](https://discuss.konghq.com/u/HYUN_SUK_JUNG)\
**Post date:** [January 23, 2020, 1:17am UTC](https://discuss.konghq.com/t/kong-1-4-k8s-db-less-504-gateway-timeout/4956/20 "2020-01-23T01:17:15Z")

</div>

It is different from my issue.  
I use cookie to save jwt token.  
If there are 10 more JTW token, Kong occur error “DNS resolution failed: dns server error: 100 cache only lookup failed”  
I saw the your error message at Nov '19.  
So I think this is the same issue but I think original issue is different
