Skip to content

out_gcs: add upload_on_shutdown to upload buffered files on exit - #12534

Open
amirgo1 wants to merge 2 commits into
fluent:masterfrom
amirgo1:out-gcs-upload-on-shutdown
Open

amirgo1 wants to merge 2 commits into
fluent:masterfrom
amirgo1:out-gcs-upload-on-shutdown

Conversation

@amirgo1

@amirgo1 amirgo1 commented Oct 8, 2026 •

Copy link
Copy Markdown

out_gcs buffers records in store_dir and uploads a file when it reaches total_file_size or upload_timeout. On exit, whatever is still buffered stays in store_dir for the next start. When the next start never comes on that disk, for example when a Kubernetes node is removed, up to upload_timeout of data per tag is lost.

This PR adds two options:

  • upload_on_shutdown (default false): upload every buffered file when Fluent Bit stops, instead of leaving it in store_dir for the next start.
  • upload_on_shutdown_timeout (default 20s, 0 means no limit): the maximum time spent on those uploads. Files not uploaded in time stay in store_dir.

The uploads run from cb_exit, on the pipeline thread after the output workers have been joined, where the upstream is in sync mode and blocking requests are safe. Each request's connect and IO timeouts are clamped to the remaining budget. A file that fails to upload stays in store_dir, as today; with preserve_data_ordering the step stops at the first failure. Nothing is uploaded on hot reload, where the new instance recovers store_dir right away. The upload logic of process_upload_queue() is factored into upload_queue_entry() and reused.

This PR also moves the upload timer to cb_worker_init. When cb_init found a backlog in store_dir, it created the timer on the engine scheduler and uploaded the backlog synchronously during init, so every later upload ran on the pipeline thread and a large backlog delayed startup. With workers, the timer is now created on the worker and the backlog is uploaded from a one-shot timer shortly after start. Without workers the behavior is unchanged.


Enter [N/A] in the box, if an item is not applicable to your change.

Testing
Before we can approve your change; please submit the following in a comment:

  • Example configuration file for the change
  • Debug log output from testing the change
  • Attached Valgrind output that shows no leaks or memory corruption was found

If this is a change to packaging of containers or native binaries then please confirm it works for all targets.

  • [N/A] Run local packaging test showing all targets (including any new ones) build.
  • [N/A] Set ok-package-test label to test for all targets (requires maintainer to do).

Documentation

  • Documentation required for this feature

fluent/fluent-bit-docs#2767

Backporting

  • [N/A] Backport to latest stable release.

Fluent Bit is licensed under Apache 2.0, by submitting this pull request I understand that this code will be released under the terms of that license.

out_gcs leaves buffered files in store_dir for the next start. When the
next start never comes, for example when a Kubernetes node is removed,
up to upload_timeout of buffered data per tag is lost.

Add upload_on_shutdown (default false) and upload_on_shutdown_timeout
(default 20s). On exit, every buffered file is sealed and uploaded from
cb_exit, which runs on the pipeline thread after the output workers have
been joined; the upstream stays in sync mode, so blocking requests are
safe there. Per-request timeouts are clamped to the remaining budget, and
files not uploaded in time stay in store_dir as before. Nothing is
uploaded on hot reload, where the new instance recovers store_dir.

Also create the upload timer in cb_worker_init. A backlog found at
startup used to create the timer from cb_init on the engine scheduler,
so every later upload ran on the pipeline thread, and the backlog was
uploaded synchronously during init. With workers, the backlog is now
uploaded from a one-shot timer on the worker shortly after start.

Signed-off-by: amirgo1 <amir.gorodetzky@wiz.io>
Signed-off-by: amirgo1 <amir.gorodetzky@wiz.io>
@coderabbitai

coderabbitai Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: defaults
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: de150099-124f-426a-81a8-87156205bf79
📥 Commits

Reviewing files that changed from the base of the PR and between d9ecd56 and 85254f4.

📒 Files selected for processing (3)
  • plugins/out_gcs/gcs.c
  • plugins/out_gcs/gcs.h
  • tests/runtime/out_gcs.c

Included review availability: This review used your included allowance. Your plan provides up to 8 included reviews per hour; 6 remain after this review.


📝 Walkthrough

Walkthrough

The GCS output plugin now supports configurable uploads during normal shutdown. It adds worker-aware upload queue initialization, retry handling, timeout validation, and tests for successful uploads, failures, timeouts, and multiple workers.

Changes

GCS shutdown uploads

Layer / File(s) Summary
Upload queue and worker startup
plugins/out_gcs/gcs.c
Queue entry processing handles uploads, retains failed files, and schedules retries. Worker initialization creates the periodic timer and schedules an upload when the queue has a backlog.
Shutdown upload configuration and execution
plugins/out_gcs/gcs.c, plugins/out_gcs/gcs.h, tests/runtime/out_gcs.c
The plugin adds shutdown-upload options and validates the timeout. Normal shutdown attempts queued uploads until the queue is empty, the timeout expires, or preserve-ordering mode encounters a failure. Tests cover upload success, recovery after failure, timeout, and multiple workers.

Priority: ⬇️ Low

Estimated code review effort: 3 (Moderate) | ~20 minutes

Change: Feature

Sequence Diagram(s)

sequenceDiagram
  participant Exit as cb_gcs_exit
  participant Store as store_dir
  participant Queue as Upload queue
  participant GCS
  Exit->>Store: Queue and seal stored files
  Exit->>Queue: Attempt uploads
  Queue->>GCS: Send upload request
  GCS-->>Queue: Return upload result
  Queue->>Store: Delete successfully uploaded file
Loading

Merge Risk: 🔵 Low · up to 85254

Shutdown may take longer than the configured upload timeout when authentication or network I/O is slow. Clarify whether the limit is best-effort or enforce it across the full upload before merging.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 30.00% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 20 functions across 3 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely describes the primary change: adding upload_on_shutdown to upload buffered files when out_gcs exits.
  • Fix all pre-merge checks with AI
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant