# How to create an Extractor that would store info in specific fields without changing the original data

**URL:** <https://community.graylog.org/t/how-to-create-an-extractor-that-would-store-info-in-specific-fields-without-changing-the-original-data/31339>\
**Category:** Graylog Central (peer support)\
**Created:** [January 26, 2024, 1:58pm UTC](https://community.graylog.org/t/how-to-create-an-extractor-that-would-store-info-in-specific-fields-without-changing-the-original-data/31339 "2024-01-26T13:58:31Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![fvr\_flho](https://avatars.discourse-cdn.com/v4/letter/f/e47c2d/32.png) [@fvr\_flho](https://community.graylog.org/u/fvr_flho)\
**Post date:** [January 26, 2024, 1:58pm UTC](https://community.graylog.org/t/how-to-create-an-extractor-that-would-store-info-in-specific-fields-without-changing-the-original-data/31339/1 "2024-01-26T13:58:31Z")

</div>

Hi

Partially inspired by LogPoint - I would like to create a series of extractors, that would ensure a uniform search to be sufficient on logs, no matter the source - eg. LogPoint will add “device\_ip=”, “device\_name=” etc. (will add a picture).

 ![logpoint](https://us1.discourse-cdn.com/flex016/uploads/graylog/original/3X/1/6/167ddb977f309da19303483ea3a6f8cad3ba7d47.png)

I know that would need a series of extractors for different logsources… but is it possible? realistic? and how?

Best regards  
/Flemming

---

<div class="post-metadata">

**Author:** ![Joel\_Duffield](https://avatars.discourse-cdn.com/v4/letter/j/71c47a/32.png) [@Joel\_Duffield](https://community.graylog.org/u/Joel_Duffield)\
**Post date:** [January 27, 2024, 4:40pm UTC](https://community.graylog.org/t/how-to-create-an-extractor-that-would-store-info-in-specific-fields-without-changing-the-original-data/31339/2 "2024-01-27T16:40:39Z")

</div>

How does the message field actually look when it arrives in Graylog, the top example is key value pairs and the bottom in JSON. Either could be dealt with very easily in a pipeline rules, and it will just create a field for each key with the value in it, these structured logs are the best kind because they are so easy to parse.

---

<div class="post-metadata">

**Author:** ![fvr\_flho](https://avatars.discourse-cdn.com/v4/letter/f/e47c2d/32.png) [@fvr\_flho](https://community.graylog.org/u/fvr_flho)\
**Post date:** [January 29, 2024, 8:15am UTC](https://community.graylog.org/t/how-to-create-an-extractor-that-would-store-info-in-specific-fields-without-changing-the-original-data/31339/3 "2024-01-29T08:15:50Z")

</div>

Hi Joel

This message example is just a screenshot found online… the pipeline rules you are talking about, could we make that so the structured fields are added? so the original source log is unchanged? (yes… I know… I’m a noobie 🙂 )

Best regards  
/Flemming

---

<div class="post-metadata">

**Author:** ![Joel\_Duffield](https://avatars.discourse-cdn.com/v4/letter/j/71c47a/32.png) [@Joel\_Duffield](https://community.graylog.org/u/Joel_Duffield)\
**Post date:** [January 29, 2024, 1:40pm UTC](https://community.graylog.org/t/how-to-create-an-extractor-that-would-store-info-in-specific-fields-without-changing-the-original-data/31339/4 "2024-01-29T13:40:09Z")

</div>

Yes, generally when you do this kind of thing you will leave the message field alone, then you may need to have a pipeline rule with some regex that clean up the format first, often there is some text before the JSON or things like that. Then you can either use key values or flatten\_json to, then lastly you will write that back into the message using set\_fields, and you will have a whole bunch of new fields with corresponding values in it.

---

<div class="post-metadata">

**Author:** ![frantz](https://avatars.discourse-cdn.com/v4/letter/f/dbc845/32.png) [@frantz](https://community.graylog.org/u/frantz)\
**Post date:** [February 2, 2024, 10:07am UTC](https://community.graylog.org/t/how-to-create-an-extractor-that-would-store-info-in-specific-fields-without-changing-the-original-data/31339/5 "2024-02-02T10:07:47Z")

</div>

The equivalent in Graylog for device\_name and device\_ip are probably source and gl2\_remote\_ip (depeding how you collect logs).

Then you need extractors or pipeline rules to extract more fields.  
If you want to automatically resolve IP/hostname you can create a lookup and create a pipeline rule or extractor that uses the lookup

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/flex016/uploads/graylog/original/3X/c/7/c7c09c6b5099570133d6502b83f50ba4430de5b6.png) [@system](https://community.graylog.org/u/system)\
**Post date:** [February 16, 2024, 10:08am UTC](https://community.graylog.org/t/how-to-create-an-extractor-that-would-store-info-in-specific-fields-without-changing-the-original-data/31339/6 "2024-02-16T10:08:39Z")

</div>

This topic was automatically closed 14 days after the last reply. New replies are no longer allowed.
