Fluentd is an open source log collector that reads events from many sources, transforms them and routes them to one or more destinations through a plugin system. Every event carries a tag, and the configuration decides which filters and outputs apply to which tags. In this tutorial you will install Fluentd on Ubuntu 24.04 using fluent-package (the successor of the old td-agent), build a pipeline that parses Nginx access logs, enriches and filters them, writes them to local files and optionally to Elasticsearch, and configure a persistent buffer so no data is lost during outages.

Prerequisites

To follow this tutorial you need:

  • A server running Ubuntu 24.04 LTS with at least 1 GB of RAM, for example a CubePath VPS.
  • A non-root user with sudo privileges.
  • Nginx installed and serving traffic (sudo apt install nginx), used as the example log source.
  • Optionally, an Elasticsearch or OpenSearch cluster reachable from the server for Step 6.

Step 1 - Installing fluent-package

fluent-package bundles Fluentd with its own Ruby runtime and a set of commonly used plugins, so it does not depend on the system Ruby. The project provides an install script that adds its APT repository and installs the package. Download the script for the long-term support (LTS) release, review it, then run it:

curl -fsSL -o install-fluent.sh https://fluentd.cdn.cncf.io/sh/install-ubuntu-noble-fluent-package6-lts.sh
less install-fluent.sh
sh install-fluent.sh

The script uses sudo internally, so you may be asked for your password. When it finishes, start the service and enable it at boot:

sudo systemctl enable --now fluentd
systemctl status fluentd --no-pager

The output should include Active: active (running). Check the version:

/opt/fluent/bin/fluentd --version
fluentd 1.19.0

The important locations are:

ItemPath
Main configuration/etc/fluent/fluentd.conf
Fluentd's own log/var/log/fluent/fluentd.log
Binaries/opt/fluent/bin/
Servicefluentd.service, running as the _fluentd user

Step 2 - Understanding sources, filters and matches

A Fluentd configuration is made of three kinds of directives:

  • <source> defines an input and assigns a tag to each event, for example nginx.access.
  • <filter PATTERN> modifies or drops events whose tag matches PATTERN.
  • <match PATTERN> sends events to an output. The first <match> that fits wins, so order matters.

Patterns use * for one tag part and ** for zero or more parts: nginx.* matches nginx.access but not nginx.access.gz, while nginx.** matches both. Filters run in the order they appear, before the matching <match>.

Step 3 - Building a pipeline for Nginx access logs

Nginx logs belong to the adm group on Ubuntu, and Fluentd runs as _fluentd. Add the user to that group so it can read them:

sudo usermod -aG adm _fluentd

Back up the packaged configuration and open the file:

sudo cp /etc/fluent/fluentd.conf /etc/fluent/fluentd.conf.orig
sudo nano /etc/fluent/fluentd.conf

Replace its content with the following pipeline:

# Test input: accepts events over HTTP on localhost only
<source>
  @type http
  bind 127.0.0.1
  port 9880
</source>

# Tail the Nginx access log
<source>
  @type tail
  path /var/log/nginx/access.log
  pos_file /var/log/fluent/nginx-access.pos
  tag nginx.access
  <parse>
    @type nginx
  </parse>
</source>

# Drop health checks
<filter nginx.access>
  @type grep
  <exclude>
    key path
    pattern /^\/(health|metrics)/
  </exclude>
</filter>

# Add the hostname to every event
<filter nginx.**>
  @type record_transformer
  <record>
    hostname "#{Socket.gethostname}"
    service nginx
  </record>
</filter>

# Events posted to the test input are printed to Fluentd's log
<match debug.**>
  @type stdout
</match>

# Nginx events: print them and write them to hourly files
<match nginx.**>
  @type copy
  <store>
    @type stdout
  </store>
  <store>
    @type file
    path /var/log/fluent/nginx/access
    <buffer time>
      timekey 1h
      timekey_wait 5m
    </buffer>
  </store>
</match>

What this does:

  • The tail source remembers its position in pos_file, so a restart does not re-send old lines. The built-in nginx parser splits the default combined format into fields like remote, method, path, code and agent.
  • The grep filter drops /health and /metrics requests; the record_transformer filter adds hostname and service. The string inside "#{...}" is evaluated once, when Fluentd starts.
  • The copy output sends every event to two stores. stdout is only there so you can see events during setup; you will remove it later. The file output groups events by hour and writes each hour to a file once the hour is over plus timekey_wait.

Validate the configuration before restarting:

sudo /opt/fluent/bin/fluentd --dry-run -c /etc/fluent/fluentd.conf

The command ends with finished dry run mode if the file is valid. Restart the service so it picks up the new configuration and the new group membership:

sudo systemctl restart fluentd

Step 4 - Verifying the pipeline

Send a test event to the HTTP input. The tag is taken from the URL path:

curl -X POST -H "Content-Type: application/json" -d '{"message":"hello fluentd"}' http://127.0.0.1:9880/debug.test

Generate some Nginx traffic, including a health check that should be filtered out:

curl -s -o /dev/null http://localhost/
curl -s -o /dev/null http://localhost/health

Now read Fluentd's log:

sudo tail -n 5 /var/log/fluent/fluentd.log
2026-09-25 10:40:12.114803516 +0000 debug.test: {"message":"hello fluentd"}
2026-09-25 10:40:20.000000000 +0000 nginx.access: {"remote":"127.0.0.1","host":"-","user":"-","method":"GET","path":"/","code":200,"size":615,"referer":"-","agent":"curl/8.5.0","hostname":"web-01","service":"nginx"}

The /health request does not appear, which confirms the grep filter works. After the current hour closes (plus five minutes), the file output creates a file in /var/log/fluent/nginx/:

ls /var/log/fluent/nginx/

Once you have confirmed events flow, remove the <store> @type stdout </store> block from the nginx.** match so Fluentd's own log does not grow with every request.

Step 5 - Installing plugins with fluent-gem

fluent-package includes the most common plugins. List the ones installed:

sudo fluent-gem list | grep '^fluent-plugin'

Any other plugin from the Fluentd plugin directory is installed with fluent-gem, which uses the bundled Ruby. For example, fluent-plugin-concat joins multi-line messages such as stack traces into one event:

sudo fluent-gem install fluent-plugin-concat
sudo systemctl restart fluentd

Do not use the system gem command: it installs into the system Ruby, which Fluentd does not use. Plugins that compile native extensions need build-essential installed first.

Step 6 - Sending events to Elasticsearch with a persistent buffer (optional)

Outputs that talk to a remote system should use a file buffer. Events are written to disk in chunks and flushed in the background; if the destination is down, chunks stay on disk and are retried, and they survive a Fluentd restart. A memory buffer is faster but loses its content if the process stops.

Check that the Elasticsearch plugin is present, and install it if it is not:

sudo fluent-gem list | grep fluent-plugin-elasticsearch || sudo fluent-gem install fluent-plugin-elasticsearch

Copy your cluster's CA certificate to /etc/fluent/es-ca.crt and make it readable by Fluentd:

sudo cp http_ca.crt /etc/fluent/es-ca.crt
sudo chmod 644 /etc/fluent/es-ca.crt

In /etc/fluent/fluentd.conf, add a third <store> inside the nginx.** match:

  <store>
    @type elasticsearch
    host your_es_host
    port 9200
    scheme https
    ca_file /etc/fluent/es-ca.crt
    user fluentd_writer
    password your_strong_password
    logstash_format true
    logstash_prefix nginx
    <buffer>
      @type file
      path /var/log/fluent/buffer/elasticsearch
      flush_interval 5s
      chunk_limit_size 8MB
      total_limit_size 1GB
      overflow_action block
      retry_max_interval 30s
      retry_forever true
    </buffer>
  </store>

Replace your_es_host, the user and your_strong_password with an Elasticsearch user that can only write nginx-* indices. The buffer settings mean:

  • flush_interval 5s: send accumulated events every five seconds.
  • total_limit_size 1GB with overflow_action block: if Elasticsearch is unreachable long enough to fill 1 GB on disk, Fluentd stops reading new input instead of discarding data. tail then resumes from pos_file when there is room.
  • retry_forever true with retry_max_interval 30s: keep retrying with exponential backoff capped at 30 seconds.

Validate and restart:

sudo /opt/fluent/bin/fluentd --dry-run -c /etc/fluent/fluentd.conf
sudo systemctl restart fluentd

If the output cannot connect, /var/log/fluent/fluentd.log shows failed to flush the buffer warnings with the retry schedule, and chunk files accumulate in /var/log/fluent/buffer/elasticsearch.

Troubleshooting

permission denied on /var/log/nginx/access.log. The _fluentd user is not in the adm group, or Fluentd was not restarted after adding it. Check with id _fluentd.

Events do not appear and there is no error. A <match> earlier in the file is catching the tag first. Move more specific matches above broad ones such as <match **>.

Unknown output plugin or Unknown filter plugin. The plugin is not installed in Fluentd's Ruby. Install it with sudo fluent-gem install fluent-plugin-NAME and restart.

Nginx lines produce pattern not matched warnings. You changed Nginx's log_format. Either switch the Nginx log to JSON and use @type json in <parse>, or write a matching @type regexp expression.

Conclusion

Fluentd is now installed from fluent-package on Ubuntu 24.04, tails and parses Nginx access logs, filters and enriches them, writes hourly files and can ship to Elasticsearch through a disk buffer that survives outages. From here you can add more sources with their own tags, run a central Fluentd aggregator that receives events from other servers with the forward input, or use the lighter Fluent Bit on each server as the forwarder and keep Fluentd for heavier processing.