# Weird Issues with two-node Cluster

**URL:** <https://community.graylog.org/t/weird-issues-with-two-node-cluster/1760>\
**Category:** Graylog Central (peer support)\
**Created:** [July 17, 2017, 6:55pm UTC](https://community.graylog.org/t/weird-issues-with-two-node-cluster/1760 "2017-07-17T18:55:19Z")\
**Posts on this page:** 19\
**Page:** 1

<div class="post-metadata">

**Author:** ![lorenzo.henriquez](https://avatars.discourse-cdn.com/v4/letter/l/90db22/32.png) [@lorenzo.henriquez](https://community.graylog.org/u/lorenzo.henriquez)\
**Post date:** [July 17, 2017, 6:55pm UTC](https://community.graylog.org/t/weird-issues-with-two-node-cluster/1760/1 "2017-07-17T18:55:19Z")

</div>

Hello everyone. I set up a two-node cluster by following the documentation [here](http://docs.graylog.org/en/2.2/pages/configuration/graylog_ctl.html#multi-vm-setup). In addition to those steps I enabled elasticsearch on the primary node as well. So in summary I have one node running all services (ES, MongoDB, Graylog, etc…) and another node only running ES.

When I go to the kopf plugin site it does show both ES nodes and all shards are assigned. However, I’m having this problem now:

- When i go to search messages and select a range of, let’s say the last hour, no messages show. But if I choose last 8 hours, all messages will show, including those coming in in real-time. By the way, this varies. Sometimes choosing messages in the last five minutes WILL show the messages correctly, and sometimes it won’t.

- Under System/Nodes I only see my master node’s information. The secondary node (running only ES) does display but no other information is shown (i.e., no memory heap usage info). And sometimes the second node doesn’t display at all.

- Finally, and I don’t know if this is normal and I’m just noticing now, but I see that many messages are marked as not processed. My journal never gets filled up though so I think that’s good (usually grows to 8% tops, from what I’ve seen but no more than that).

Please let me know what other information I should provide to make my problem more clear to you (logs, screenshots, whatever).

---

<div class="post-metadata">

**Author:** ![jochen](https://sea2.discourse-cdn.com/flex016/user_avatar/community.graylog.org/jochen/32/8_2.png) [@jochen](https://community.graylog.org/u/jochen)\
**Post date:** [July 18, 2017, 7:43am UTC](https://community.graylog.org/t/weird-issues-with-two-node-cluster/1760/2 "2017-07-18T07:43:17Z")

</div>

> [@lorenzo.henriquez](#):
>
> When i go to search messages and select a range of, let’s say the last hour, no messages show. But if I choose last 8 hours, all messages will show, including those coming in in real-time.

Do the messages have the correct timestamps? What do the index details on the _System_ / _Indices_ / _Index Set_ page say?

> [@lorenzo.henriquez](#):
>
> Under System/Nodes I only see my master node’s information.

Yes, because that page only displays the details of _Graylog_ nodes, not those of Elasticsearch nodes.

> [@lorenzo.henriquez](#):
>
> Finally, and I don’t know if this is normal and I’m just noticing now, but I see that many messages are marked as not processed.

Is there anything unusual in the logs of your Graylog and Elasticsearch nodes?  
➡ [http://docs.graylog.org/en/2.2/pages/configuration/file\_location.html#omnibus-package](http://docs.graylog.org/en/2.2/pages/configuration/file_location.html#omnibus-package)

---

<div class="post-metadata">

**Author:** ![lorenzo.henriquez](https://avatars.discourse-cdn.com/v4/letter/l/90db22/32.png) [@lorenzo.henriquez](https://community.graylog.org/u/lorenzo.henriquez)\
**Post date:** [July 18, 2017, 1:28pm UTC](https://community.graylog.org/t/weird-issues-with-two-node-cluster/1760/3 "2017-07-18T13:28:34Z")

</div>

> [@jochen](#):
>
> Do the messages have the correct timestamps? What do the index details on the System / Indices / Index Set page say?

They do have the correct timestamp.  
These are the details for the current index:

 ![](https://us1.discourse-cdn.com/flex016/uploads/graylog/original/1X/3e5c5916a45ae0e7eeefeb10e9e9027df428f8c7.png)

> [@jochen](#):
>
> Yes, because that page only displays the details of Graylog nodes, not those of Elasticsearch nodes.

Thank you for pointing that out!

> [@jochen](#):
>
> Is there anything unusual in the logs of your Graylog and Elasticsearch nodes?

I have a few warnings like this in the Graylog server (which is also a slave ES node):

```
2017-07-18_13:05:35.45455 [2017-07-18 09:05:35,452][WARN][env] [Derrick Slegers Speed] max file descriptors [64000] for elasticsearch process likely too low, consider increasing to at least [65536]

```

Also have a few like these in both ES nodes:  
\> 2017-07-18\_13:05:05.29675 [2017-07-18 09:05:05,296][INFO][cluster.service] [Veil] removed {{graylog-e12c96d2-6edc-4982-b476-a82a4238ce3c}{kjUZKScOQiOGPWFER9TD9g}{secondary\_ES\_NODE\_IP}{secondary\_ES\_NODE\_IP:9\> 350}{client=true, data=false, master=false},}, reason: zen-disco-node\_failed({graylog-e12c96d2-6edc-4982-b476-a82a4238ce3c}{kjUZKScOQiOGPWFER9TD9g}{secondary\_ES\_NODE\_IP}{secondary\_ES\_NODE\_IP:9350}{client=true, data=fa  
\> lse, master=false}), reason transport disconnected

Also noticed this from the kopf plugin page. It says that there are 3 nodes when I know for a fact there’s only two.

So when I went to the NODES tab I saw this. Notice that the first two nodes listed have the same IP although different ports. They also seem to be running different ES versions. No idea how that happened and why it’s listing two different ES nodes under the same IP:

 ![](https://us1.discourse-cdn.com/flex016/uploads/graylog/original/1X/e01013ce05f8984cd5db7710e76a7cd1fad6a07b.png)

Thank you for reply and I hope what I posted is useful.

---

<div class="post-metadata">

**Author:** ![jochen](https://sea2.discourse-cdn.com/flex016/user_avatar/community.graylog.org/jochen/32/8_2.png) [@jochen](https://community.graylog.org/u/jochen)\
**Post date:** [July 18, 2017, 2:12pm UTC](https://community.graylog.org/t/weird-issues-with-two-node-cluster/1760/5 "2017-07-18T14:12:52Z")

</div>

> [@lorenzo.henriquez](#):
>
> They do have the correct timestamp.

Are you sure about that? Do you have some examples?

How exactly do you send the messages to Graylog and how is the input receiving the messages configured?

> [@lorenzo.henriquez](#):
>
> Also noticed this from the kopf plugin page. It says that there are 3 nodes when I know for a fact there’s only two.

Graylog joins the Elasticsearch cluster as a client node (i. e. not master eligible, doesn’t store data). That’s what you see in the Elasticsearch cluster state with Kopf (or any other UI).  
You can identify the Graylog client node with the “graylog” prefix followed by the Graylog node ID.

---

<div class="post-metadata">

**Author:** ![lorenzo.henriquez](https://avatars.discourse-cdn.com/v4/letter/l/90db22/32.png) [@lorenzo.henriquez](https://community.graylog.org/u/lorenzo.henriquez)\
**Post date:** [July 18, 2017, 2:26pm UTC](https://community.graylog.org/t/weird-issues-with-two-node-cluster/1760/6 "2017-07-18T14:26:33Z")

</div>

> [@jochen](#):
>
> Are you sure about that? Do you have some examples?

Here’s a screenshot I just took now. I have highlighted both the original message’s timestamp and the one that Graylog puts in it:

 ![](https://us1.discourse-cdn.com/flex016/uploads/graylog/original/1X/6b9dcfe4f7a47949ee2f49d301b9881bf66a5656.png)

> [@jochen](#):
>
> How exactly do you send the messages to Graylog and how is the input receiving the messages configured?

For that particular message (coming from a Cisco ASA) this is the configuration in the ASA:

![](https://us1.discourse-cdn.com/flex016/uploads/graylog/original/1X/82f301d7f67d8b97bbf9f80ab339d519cc3a6f57.png)

And this is the UDP input configuration. Also, I’m using that one input for many other network devices that use syslog:

 ![](https://us1.discourse-cdn.com/flex016/uploads/graylog/original/1X/a723c25414a267a65ebf55145d95e6007a46a5ee.png)

> [@jochen](#):
>
> Graylog joins the Elasticsearch cluster as a client node (i. e. not master eligible, doesn’t store data). That’s what you see in the Elasticsearch cluster state with Kopf (or any other UI).  
> You can identify the Graylog client node with the “graylog” prefix followed by the Graylog node ID.

Thank you for elaborating on that. I still don’t quite understand why it lists the Graylog server as using a different ES version.

---

<div class="post-metadata">

**Author:** ![jochen](https://sea2.discourse-cdn.com/flex016/user_avatar/community.graylog.org/jochen/32/8_2.png) [@jochen](https://community.graylog.org/u/jochen)\
**Post date:** [July 18, 2017, 2:38pm UTC](https://community.graylog.org/t/weird-issues-with-two-node-cluster/1760/7 "2017-07-18T14:38:36Z")

</div>

> [@lorenzo.henriquez](#):
>
> Here’s a screenshot I just took now. I have highlighted both the original message’s timestamp and the one that Graylog puts in it

That looks correct (assuming that your Cisco ASA is configured to use timezone UTC).

Could you please elaborate on the issue you have (or had)?

> [@lorenzo.henriquez](#):
>
> I still don’t quite understand why it lists the Graylog server as using a different ES version.

Because that’s the version of Elasticsearch embedded in the version of Graylog you’re using.

---

<div class="post-metadata">

**Author:** ![lorenzo.henriquez](https://avatars.discourse-cdn.com/v4/letter/l/90db22/32.png) [@lorenzo.henriquez](https://community.graylog.org/u/lorenzo.henriquez)\
**Post date:** [July 18, 2017, 2:56pm UTC](https://community.graylog.org/t/weird-issues-with-two-node-cluster/1760/8 "2017-07-18T14:56:43Z")

</div>

> [@jochen](#):
>
> Could you please elaborate on the issue you have (or had)?

Ok, let’s take the screenshot with the message above as an example:

 ![](https://us1.discourse-cdn.com/flex016/uploads/graylog/original/1X/3b6819c04232d73b41fe0dda0abcd9447e995933.png)

Notice how I selected to search in the last 2 hours? If I had selected to search in the last 5 minutes, the results would’ve been no messages at all (even though the message shown happened in real-time). This issue happens randomly but it’s _almost_ constant. Moreover, at times I have to search messages in the last day in order to see messages coming in real time.

Never had this issue until I joined the new ES node. When it was just one node, the search function worked fine.

Please let me know if you need any more specific details and thanks again for your help.

---

<div class="post-metadata">

**Author:** ![jochen](https://sea2.discourse-cdn.com/flex016/user_avatar/community.graylog.org/jochen/32/8_2.png) [@jochen](https://community.graylog.org/u/jochen)\
**Post date:** [July 18, 2017, 3:01pm UTC](https://community.graylog.org/t/weird-issues-with-two-node-cluster/1760/9 "2017-07-18T15:01:17Z")

</div>

Which timezone is the system running Graylog on?  
Which timezone is the system running Elasticsearch on?  
Which timezone is configured for the Graylog user you’re logged in with?  
Which timezone is configured on your Cisco ASA devices?

---

<div class="post-metadata">

**Author:** ![lorenzo.henriquez](https://avatars.discourse-cdn.com/v4/letter/l/90db22/32.png) [@lorenzo.henriquez](https://community.graylog.org/u/lorenzo.henriquez)\
**Post date:** [July 18, 2017, 4:32pm UTC](https://community.graylog.org/t/weird-issues-with-two-node-cluster/1760/10 "2017-07-18T16:32:07Z")

</div>

I believe they’re all in the same timezone but I’ll post screenshots, just in case.

> [@jochen](#):
>
> Which timezone is the system running Graylog on?

![](https://us1.discourse-cdn.com/flex016/uploads/graylog/original/1X/e688b12d87d6261250aefb74379763af04941fe9.png)

> [@jochen](#):
>
> Which timezone is the system running Elasticsearch on?

![](https://us1.discourse-cdn.com/flex016/uploads/graylog/original/1X/96f2877e05da955ff455555c27b004b08b5e6206.png)

> [@jochen](#):
>
> Which timezone is configured for the Graylog user you’re logged in with?

 ![](https://us1.discourse-cdn.com/flex016/uploads/graylog/original/1X/01435304ab350a71fda149388ea2429c25290f5c.png)

> [@jochen](#):
>
> Which timezone is configured on your Cisco ASA devices?

![](https://us1.discourse-cdn.com/flex016/uploads/graylog/original/1X/69ccec662a68b3a981c25531e59eca63ef5ce501.png)

---

<div class="post-metadata">

**Author:** ![lorenzo.henriquez](https://avatars.discourse-cdn.com/v4/letter/l/90db22/32.png) [@lorenzo.henriquez](https://community.graylog.org/u/lorenzo.henriquez)\
**Post date:** [July 18, 2017, 5:19pm UTC](https://community.graylog.org/t/weird-issues-with-two-node-cluster/1760/11 "2017-07-18T17:19:21Z")

</div>

UPDATE: Just found out that recalculating the indexes manually (Indices\>Maintenance\>Recalculate index ranges) fixes the problem temporarily. It fails again in exactly 5 minutes after the indexing task is finished (because I’m searching messages in the last 5 minutes of course). If I choose to search in the last 15 minutes, then messages will show…until 15 minutes from the last indexing have passed.

Any thoughts?

---

<div class="post-metadata">

**Author:** ![jtkarvo](https://avatars.discourse-cdn.com/v4/letter/j/43a26b/32.png) [@jtkarvo](https://community.graylog.org/u/jtkarvo)\
**Post date:** [July 18, 2017, 5:25pm UTC](https://community.graylog.org/t/weird-issues-with-two-node-cluster/1760/12 "2017-07-18T17:25:11Z")

</div>

At least you should increase the number of file descriptors, as the logs suggest. Those broken connections may well result from that reason. You could use something much larger, such as 260000.

Also, check using the kopf, what is the memory usage in the ES nodes; just to be sure.

---

<div class="post-metadata">

**Author:** ![lorenzo.henriquez](https://avatars.discourse-cdn.com/v4/letter/l/90db22/32.png) [@lorenzo.henriquez](https://community.graylog.org/u/lorenzo.henriquez)\
**Post date:** [July 18, 2017, 5:27pm UTC](https://community.graylog.org/t/weird-issues-with-two-node-cluster/1760/13 "2017-07-18T17:27:23Z")

</div>

> [@jtkarvo](#):
>
> Also, check using the kopf, what is the memory usage in the ES nodes; just to be sure.

![](https://us1.discourse-cdn.com/flex016/uploads/graylog/original/1X/413d0f05ce5b5108b463a4e7d697d90d73f97792.png)

Also, I’ll look into increasing the number of file descriptors.

---

<div class="post-metadata">

**Author:** ![lorenzo.henriquez](https://avatars.discourse-cdn.com/v4/letter/l/90db22/32.png) [@lorenzo.henriquez](https://community.graylog.org/u/lorenzo.henriquez)\
**Post date:** [July 18, 2017, 5:43pm UTC](https://community.graylog.org/t/weird-issues-with-two-node-cluster/1760/14 "2017-07-18T17:43:17Z")

</div>

I found out someone having the same issue [here](https://groups.google.com/forum/#!msg/graylog2/BRV9FgQbz1A/-JqrfiCSAgAJ). For one of the people in the thread, manually cycling the deflector fixed their issues. Do you know what unexpected consequences that may have?

---

<div class="post-metadata">

**Author:** ![jtkarvo](https://avatars.discourse-cdn.com/v4/letter/j/43a26b/32.png) [@jtkarvo](https://community.graylog.org/u/jtkarvo)\
**Post date:** [July 18, 2017, 5:53pm UTC](https://community.graylog.org/t/weird-issues-with-two-node-cluster/1760/15 "2017-07-18T17:53:45Z")

</div>

hi,

if you have set up retention based on number of indices, you might lose data sooner than you intended. You can counter this by having a few indices spare in the retention settings, so that you have the possibility to manually cycle the deflector index.

---

<div class="post-metadata">

**Author:** ![lorenzo.henriquez](https://avatars.discourse-cdn.com/v4/letter/l/90db22/32.png) [@lorenzo.henriquez](https://community.graylog.org/u/lorenzo.henriquez)\
**Post date:** [July 18, 2017, 5:57pm UTC](https://community.graylog.org/t/weird-issues-with-two-node-cluster/1760/16 "2017-07-18T17:57:25Z")

</div>

Hi and thanks for your reply. These are my current retention settings. What would suggest I change? Sorry I’m not too clear on all this. Never really had problems before with Graylog.

 ![](https://us1.discourse-cdn.com/flex016/uploads/graylog/original/1X/f7c7b1ea203db52c4b46535ef9d083ad5126369e.png)

---

<div class="post-metadata">

**Author:** ![jtkarvo](https://avatars.discourse-cdn.com/v4/letter/j/43a26b/32.png) [@jtkarvo](https://community.graylog.org/u/jtkarvo)\
**Post date:** [July 18, 2017, 6:35pm UTC](https://community.graylog.org/t/weird-issues-with-two-node-cluster/1760/17 "2017-07-18T18:35:33Z")

</div>

That depends on your policy.

For example, if your policy is to hold data for 30 days, you could use index rotation based on time, and then choose some time (such as 1 day or 4 hours or whatever is convenient for you) as retention time. Then max number of indices would be 30 days / rotation time + some spare for manual deflector cycling.

Here your policy seems to be to hold the last 800,000,000 messages. Then, to have spare, you can add some additional indices to the max number of indices, so that you have some spare for deflector cycling.

For example, if you add 3 to the max number of indices, you can manually cycle your deflector 3 times without breaking your company policy.

---

<div class="post-metadata">

**Author:** ![lorenzo.henriquez](https://avatars.discourse-cdn.com/v4/letter/l/90db22/32.png) [@lorenzo.henriquez](https://community.graylog.org/u/lorenzo.henriquez)\
**Post date:** [July 18, 2017, 6:38pm UTC](https://community.graylog.org/t/weird-issues-with-two-node-cluster/1760/18 "2017-07-18T18:38:40Z")

</div>

Got it, thanks for the explanation.

---

<div class="post-metadata">

**Author:** ![lorenzo.henriquez](https://avatars.discourse-cdn.com/v4/letter/l/90db22/32.png) [@lorenzo.henriquez](https://community.graylog.org/u/lorenzo.henriquez)\
**Post date:** [July 19, 2017, 12:55am UTC](https://community.graylog.org/t/weird-issues-with-two-node-cluster/1760/19 "2017-07-19T00:55:06Z")

</div>

FYI, I manually cycled the deflector early this afternoon and the problem seems to be gone. Thank you both for your useful insights.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/flex016/uploads/graylog/original/3X/c/7/c7c09c6b5099570133d6502b83f50ba4430de5b6.png) [@system](https://community.graylog.org/u/system)\
**Post date:** [August 2, 2017, 12:55am UTC](https://community.graylog.org/t/weird-issues-with-two-node-cluster/1760/20 "2017-08-02T00:55:14Z")

</div>

This topic was automatically closed 14 days after the last reply. New replies are no longer allowed.
