# Issues with Graylog after moving to an elasticsearch cluster

**URL:** <https://community.graylog.org/t/issues-with-graylog-after-moving-to-an-elasticsearch-cluster/5521>\
**Category:** Graylog Central (peer support)\
**Created:** [June 8, 2018, 8:37am UTC](https://community.graylog.org/t/issues-with-graylog-after-moving-to-an-elasticsearch-cluster/5521 "2018-06-08T08:37:04Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![mariusgeonea](https://sea2.discourse-cdn.com/flex016/user_avatar/community.graylog.org/mariusgeonea/32/2162_2.png) [@mariusgeonea](https://community.graylog.org/u/mariusgeonea)\
**Post date:** [June 8, 2018, 8:37am UTC](https://community.graylog.org/t/issues-with-graylog-after-moving-to-an-elasticsearch-cluster/5521/1 "2018-06-08T08:37:04Z")

</div>

Hello,

after i moved my setup from a single elasticsearch server to a cluster (in the cluster i have 3 Master nodes and 3 Data nodes) the output in the right corner decreased from 7500 msg/s to arround 4500 msg/s, and my input was the same as vefore around 7500msg/s.

what i did so far, i have checked the time and data, everywhere is the same, because it’s synced via NTP.  
i have also configured the following on elasticsearch:

```auto
# Recover only after the given number of nodes have joined the cluster. Can be seen as "minimum number of nodes to attempt recovery at all".
gateway.recover_after_nodes: 4
# Time to wait for additional nodes after recover_after_nodes is met.
gateway.recover_after_time: 5m
# Inform ElasticSearch how many nodes form a full cluster. If this number is met, start up immediately.
gateway.expected_nodes: 6

```

then i did 300mb instead of 150:

```auto
indices.store.throttle.max_bytes_per_sec: 300mb

```

in graylog i changed the process buffer to 20 from 10  
also in graylog i rotated all the indexes manually and deleted the old ones.

with all of these nothing changed in better.

have you ever seen this problem before, or any ideas that can help me?

Thanks,  
Marius.

---

<div class="post-metadata">

**Author:** ![jochen](https://sea2.discourse-cdn.com/flex016/user_avatar/community.graylog.org/jochen/32/8_2.png) [@jochen](https://community.graylog.org/u/jochen)\
**Post date:** [June 8, 2018, 8:48am UTC](https://community.graylog.org/t/issues-with-graylog-after-moving-to-an-elasticsearch-cluster/5521/2 "2018-06-08T08:48:15Z")

</div>

Which version of Graylog are you using?  
Which version of Elasticsearch are you using?  
Was the single Elasticsearch node previously running on the same machine as the Graylog node?  
What’s the complete configuration of Graylog and Elasticsearch?  
What’s in the logs of your Graylog and Elasticsearch nodes?  
➡ [http://docs.graylog.org/en/2.4/pages/configuration/file\_location.html](http://docs.graylog.org/en/2.4/pages/configuration/file_location.html)

---

<div class="post-metadata">

**Author:** ![mariusgeonea](https://sea2.discourse-cdn.com/flex016/user_avatar/community.graylog.org/mariusgeonea/32/2162_2.png) [@mariusgeonea](https://community.graylog.org/u/mariusgeonea)\
**Post date:** [June 8, 2018, 9:19am UTC](https://community.graylog.org/t/issues-with-graylog-after-moving-to-an-elasticsearch-cluster/5521/3 "2018-06-08T09:19:22Z")

</div>

Q: Which version of Graylog are you using?  
A: Graylog v2.4.5+8e18e6a

Q: Which version of Elasticsearch are you using?  
A: elasticsearch-5.6

Q: Was the single Elasticsearch node previously running on the same machine as the Graylog node?  
A: no

Q: What’s the complete configuration of Graylog and Elasticsearch please note that from the graylog conf file some things were deleted due to privacy issues?  
A: graylog - [https://drive.google.com/open?id=15nwuTCj19gJOSj0Fkv7v09ZLUkPK\_cE-](https://drive.google.com/open?id=15nwuTCj19gJOSj0Fkv7v09ZLUkPK_cE-)  
elastic master node - [https://drive.google.com/open?id=1jaCuX8hsNkY0Q5TEfBHPTMgHmlga4b9u](https://drive.google.com/open?id=1jaCuX8hsNkY0Q5TEfBHPTMgHmlga4b9u)  
data node - [https://drive.google.com/open?id=1CarnWV7sfbRzLA08wwEexdfFdd1hN6Js](https://drive.google.com/open?id=1CarnWV7sfbRzLA08wwEexdfFdd1hN6Js)

Q: What’s in the logs of your Graylog and Elasticsearch nodes?  
A: [https://drive.google.com/open?id=1E6Ftnw8aWjBDpAz2MG3NHWcOBJ21V98J](https://drive.google.com/open?id=1E6Ftnw8aWjBDpAz2MG3NHWcOBJ21V98J)

---

<div class="post-metadata">

**Author:** ![jochen](https://sea2.discourse-cdn.com/flex016/user_avatar/community.graylog.org/jochen/32/8_2.png) [@jochen](https://community.graylog.org/u/jochen)\
**Post date:** [June 8, 2018, 9:45am UTC](https://community.graylog.org/t/issues-with-graylog-after-moving-to-an-elasticsearch-cluster/5521/4 "2018-06-08T09:45:17Z")

</div>

> [@mariusgeonea](#):
>
> Q: What’s in the logs of your Graylog and Elasticsearch nodes?  
> A: [logs - Google Drive](https://drive.google.com/open?id=1E6Ftnw8aWjBDpAz2MG3NHWcOBJ21V98J)

There are no files in that folder.

---

<div class="post-metadata">

**Author:** ![mariusgeonea](https://sea2.discourse-cdn.com/flex016/user_avatar/community.graylog.org/mariusgeonea/32/2162_2.png) [@mariusgeonea](https://community.graylog.org/u/mariusgeonea)\
**Post date:** [June 8, 2018, 10:45am UTC](https://community.graylog.org/t/issues-with-graylog-after-moving-to-an-elasticsearch-cluster/5521/5 "2018-06-08T10:45:40Z")

</div>

[https://drive.google.com/file/d/1nc2OC\_56QD321J2CfUCB0IHZJSaqHK\_e/view?usp=sharing](https://drive.google.com/file/d/1nc2OC_56QD321J2CfUCB0IHZJSaqHK_e/view?usp=sharing)  
[https://drive.google.com/file/d/1G32upI1mU1Fc4CAEaWf6O483pTnRsBMc/view?usp=sharing](https://drive.google.com/file/d/1G32upI1mU1Fc4CAEaWf6O483pTnRsBMc/view?usp=sharing)  
[https://drive.google.com/file/d/1ob4ppMqRFE\_rV-TlpFJAC9jzcWeEL04K/view?usp=sharing](https://drive.google.com/file/d/1ob4ppMqRFE_rV-TlpFJAC9jzcWeEL04K/view?usp=sharing)  
[https://drive.google.com/file/d/1HKDphJKGlZkzimU9r1QSSz71PKELeVPb/view?usp=sharing](https://drive.google.com/file/d/1HKDphJKGlZkzimU9r1QSSz71PKELeVPb/view?usp=sharing)  
[https://drive.google.com/file/d/15GEbbCkR7EjxaBMwxGBUVHD-HSIzg7AB/view?usp=sharing](https://drive.google.com/file/d/15GEbbCkR7EjxaBMwxGBUVHD-HSIzg7AB/view?usp=sharing)  
[https://drive.google.com/file/d/1AGSmrimcDefuEziYfYk2hnzs1\_ZxaHnU/view?usp=sharing](https://drive.google.com/file/d/1AGSmrimcDefuEziYfYk2hnzs1_ZxaHnU/view?usp=sharing)  
[https://drive.google.com/file/d/1u9hQaa2--2zuASqrfaxswiN5\_1zdpJie/view?usp=sharing](https://drive.google.com/file/d/1u9hQaa2--2zuASqrfaxswiN5_1zdpJie/view?usp=sharing)

---

<div class="post-metadata">

**Author:** ![jochen](https://sea2.discourse-cdn.com/flex016/user_avatar/community.graylog.org/jochen/32/8_2.png) [@jochen](https://community.graylog.org/u/jochen)\
**Post date:** [June 8, 2018, 10:59am UTC](https://community.graylog.org/t/issues-with-graylog-after-moving-to-an-elasticsearch-cluster/5521/6 "2018-06-08T10:59:50Z")

</div>

> ```plaintext
> 2018-06-08T03:57:16.436-04:00 WARN [Messages] Failed to index message: index=<firewall_deflector> id=<725de9a3-6af1-11e8-82eb-0050568640e7> error=<{"type":"invalid_index_name_exception","reason":"Invalid index name [firewall_deflector], already exists as alias","index_uuid":"_na_","index":"firewall_deflector"}>
> 
> ```

[http://docs.graylog.org/en/2.4/pages/faq.html#how-do-i-fix-the-deflector-exists-as-an-index-and-is-not-an-alias-error-message](http://docs.graylog.org/en/2.4/pages/faq.html#how-do-i-fix-the-deflector-exists-as-an-index-and-is-not-an-alias-error-message)

There are a lot of warnings and error messages in these logs, which you should try to resolve individually. If the performance is still worse after that, come back and post the current logs (and configuration files if anything changed).

---

<div class="post-metadata">

**Author:** ![mariusgeonea](https://sea2.discourse-cdn.com/flex016/user_avatar/community.graylog.org/mariusgeonea/32/2162_2.png) [@mariusgeonea](https://community.graylog.org/u/mariusgeonea)\
**Post date:** [June 8, 2018, 11:30am UTC](https://community.graylog.org/t/issues-with-graylog-after-moving-to-an-elasticsearch-cluster/5521/7 "2018-06-08T11:30:00Z")

</div>

how to i resolve those? by deleting the current logs?

---

<div class="post-metadata">

**Author:** ![jochen](https://sea2.discourse-cdn.com/flex016/user_avatar/community.graylog.org/jochen/32/8_2.png) [@jochen](https://community.graylog.org/u/jochen)\
**Post date:** [June 8, 2018, 11:46am UTC](https://community.graylog.org/t/issues-with-graylog-after-moving-to-an-elasticsearch-cluster/5521/8 "2018-06-08T11:46:48Z")

</div>

Please read the linked FAQ article.

---

<div class="post-metadata">

**Author:** ![mariusgeonea](https://sea2.discourse-cdn.com/flex016/user_avatar/community.graylog.org/mariusgeonea/32/2162_2.png) [@mariusgeonea](https://community.graylog.org/u/mariusgeonea)\
**Post date:** [June 8, 2018, 12:44pm UTC](https://community.graylog.org/t/issues-with-graylog-after-moving-to-an-elasticsearch-cluster/5521/9 "2018-06-08T12:44:22Z")

</div>

no more invalid indexes

but the output messages didn’t increase…

---

<div class="post-metadata">

**Author:** ![mariusgeonea](https://sea2.discourse-cdn.com/flex016/user_avatar/community.graylog.org/mariusgeonea/32/2162_2.png) [@mariusgeonea](https://community.graylog.org/u/mariusgeonea)\
**Post date:** [June 8, 2018, 1:22pm UTC](https://community.graylog.org/t/issues-with-graylog-after-moving-to-an-elasticsearch-cluster/5521/10 "2018-06-08T13:22:33Z")

</div>

i re-did the elastic cluster from 0…same issue nothing good is happening…  
i;m really running outta ideas here…  
 ![image](https://us1.discourse-cdn.com/flex016/uploads/graylog/original/2X/5/57f839fd0485c34f9553173de06d45f5a7cd4200.png)

---

<div class="post-metadata">

**Author:** ![mariusgeonea](https://sea2.discourse-cdn.com/flex016/user_avatar/community.graylog.org/mariusgeonea/32/2162_2.png) [@mariusgeonea](https://community.graylog.org/u/mariusgeonea)\
**Post date:** [June 8, 2018, 1:36pm UTC](https://community.graylog.org/t/issues-with-graylog-after-moving-to-an-elasticsearch-cluster/5521/11 "2018-06-08T13:36:34Z")

</div>

another thing is that the CPU is in 100% all the time

---

<div class="post-metadata">

**Author:** ![jochen](https://sea2.discourse-cdn.com/flex016/user_avatar/community.graylog.org/jochen/32/8_2.png) [@jochen](https://community.graylog.org/u/jochen)\
**Post date:** [June 8, 2018, 1:40pm UTC](https://community.graylog.org/t/issues-with-graylog-after-moving-to-an-elasticsearch-cluster/5521/12 "2018-06-08T13:40:04Z")

</div>

Try playing around with the batch sizes and batch commit intervals:

> <https://github.com/Graylog2/graylog2-server/blob/2.4.5/misc/graylog.conf#L349-L359>

---

<div class="post-metadata">

**Author:** ![mariusgeonea](https://sea2.discourse-cdn.com/flex016/user_avatar/community.graylog.org/mariusgeonea/32/2162_2.png) [@mariusgeonea](https://community.graylog.org/u/mariusgeonea)\
**Post date:** [June 8, 2018, 1:53pm UTC](https://community.graylog.org/t/issues-with-graylog-after-moving-to-an-elasticsearch-cluster/5521/13 "2018-06-08T13:53:54Z")

</div>

right now i have :

```nohighlight
# Batch size for the Elasticsearch output. This is the maximum (!) number of messages the Elasticsearch output
# module will get at once and write to Elasticsearch in a batch call. If the configured batch size has not been
# reached within output_flush_interval seconds, everything that is available will be flushed at once. Remember
# that every outputbuffer processor manages its own batch and performs its own batch write calls.
# ("outputbuffer_processors" variable)
output_batch_size = 1000

# Flush interval (in seconds) for the Elasticsearch output. This is the maximum amount of time between two
# batches of messages written to Elasticsearch. It is only effective at all if your minimum number of messages
# for this time period is less than output_batch_size * outputbuffer_processors.
output_flush_interval = 2

# As stream outputs are loaded only on demand, an output which is failing to initialize will be tried over and
# over again. To prevent this, the following configuration options define after how many faults an output will
# not be tried again for an also configurable amount of seconds.
output_fault_count_threshold = 5
output_fault_penalty_seconds = 30

# The number of parallel running processors.
# Raise this number if your buffers are filling up.

processbuffer_processors = 15
outputbuffer_processors = 10

```

with the same result

---

<div class="post-metadata">

**Author:** ![jan](https://sea2.discourse-cdn.com/flex016/user_avatar/community.graylog.org/jan/32/11_2.png) [@jan](https://community.graylog.org/u/jan)\
**Post date:** [June 8, 2018, 6:47pm UTC](https://community.graylog.org/t/issues-with-graylog-after-moving-to-an-elasticsearch-cluster/5521/14 "2018-06-08T18:47:35Z")

</div>

What are your sharding and replication settings? Did you set the refresh interval for elasticsearch?

In addition I would make the check list work with the following elastic articel

[https://www.elastic.co/guide/en/elasticsearch/reference/5.6/tune-for-indexing-speed.html](https://www.elastic.co/guide/en/elasticsearch/reference/5.6/tune-for-indexing-speed.html)

I would raise the output\_batch\_size to 4000 and lower the outputbuffer\_processors to 5, I have the feeling that your updates are eating all available elasticsearch threats and you need to push a higher amount of message with lower amount of connections.

---

<div class="post-metadata">

**Author:** ![mariusgeonea](https://sea2.discourse-cdn.com/flex016/user_avatar/community.graylog.org/mariusgeonea/32/2162_2.png) [@mariusgeonea](https://community.graylog.org/u/mariusgeonea)\
**Post date:** [June 9, 2018, 3:32am UTC](https://community.graylog.org/t/issues-with-graylog-after-moving-to-an-elasticsearch-cluster/5521/15 "2018-06-09T03:32:10Z")

</div>

my shards were 4 /index and i have 6 indexes, unfortunately i didn’t test that one out…  
what i did i raised the CPU cores from 16 to 24 and now everything is running fine again with a CPU utilization of 50%. i will try to test what you have suggested to see if the CPU usage decreases.

now i’m trying to test with 2 shards /index to see if there is any change in the storage space, maybe that will also help reduce the CPU usage…

---

<div class="post-metadata">

**Author:** ![jan](https://sea2.discourse-cdn.com/flex016/user_avatar/community.graylog.org/jan/32/11_2.png) [@jan](https://community.graylog.org/u/jan)\
**Post date:** [June 9, 2018, 2:04pm UTC](https://community.graylog.org/t/issues-with-graylog-after-moving-to-an-elasticsearch-cluster/5521/16 "2018-06-09T14:04:36Z")

</div>

the total number of outbuffer, inputbuffer and processbuffer processors should be 3/4 of the available cores max.

---

<div class="post-metadata">

**Author:** ![mariusgeonea](https://sea2.discourse-cdn.com/flex016/user_avatar/community.graylog.org/mariusgeonea/32/2162_2.png) [@mariusgeonea](https://community.graylog.org/u/mariusgeonea)\
**Post date:** [June 9, 2018, 2:28pm UTC](https://community.graylog.org/t/issues-with-graylog-after-moving-to-an-elasticsearch-cluster/5521/17 "2018-06-09T14:28:30Z")

</div>

3 is the number of processors and 4 are the CPU cores?

i’m afraid i don’t really understand the formula

---

<div class="post-metadata">

**Author:** ![mariusgeonea](https://sea2.discourse-cdn.com/flex016/user_avatar/community.graylog.org/mariusgeonea/32/2162_2.png) [@mariusgeonea](https://community.graylog.org/u/mariusgeonea)\
**Post date:** [June 9, 2018, 2:42pm UTC](https://community.graylog.org/t/issues-with-graylog-after-moving-to-an-elasticsearch-cluster/5521/18 "2018-06-09T14:42:19Z")

</div>

i think i get, probably is three quarters, i hope that’s it that means 75%, so if i have 24 cores i should give it 16 in total, now i have set it to outbuffer 6, input 2 and 10 for the processbuffer.

i hope this is what you meant

---

<div class="post-metadata">

**Author:** ![jan](https://sea2.discourse-cdn.com/flex016/user_avatar/community.graylog.org/jan/32/11_2.png) [@jan](https://community.graylog.org/u/jan)\
**Post date:** [June 10, 2018, 11:30am UTC](https://community.graylog.org/t/issues-with-graylog-after-moving-to-an-elasticsearch-cluster/5521/19 "2018-06-10T11:30:18Z")

</div>

you got it - sorry that I did not wrote it more clear

---

<div class="post-metadata">

**Author:** ![mariusgeonea](https://sea2.discourse-cdn.com/flex016/user_avatar/community.graylog.org/mariusgeonea/32/2162_2.png) [@mariusgeonea](https://community.graylog.org/u/mariusgeonea)\
**Post date:** [June 10, 2018, 12:30pm UTC](https://community.graylog.org/t/issues-with-graylog-after-moving-to-an-elasticsearch-cluster/5521/20 "2018-06-10T12:30:07Z")

</div>

> [@jan](#):
>
> the total number of outbuffer, inputbuffer and processbuffer processors should be 3/4 of the available cores max.

no worries. in fact you were pretty clear, but because i’m not used to math that much i didn’t understood those three quarters 🙂

anyway, at the very moment i’m running those plus an output batch size of 12000, do you think i can go to 20000, and what could happen in the worse case if an output batch of 20k or 30k won’t work?

[Next page](https://community.graylog.org/t/issues-with-graylog-after-moving-to-an-elasticsearch-cluster/5521.md?page=2)
