Repository navigation
Conversation
out_gcs leaves buffered files in store_dir for the next start. When the next start never comes, for example when a Kubernetes node is removed, up to upload_timeout of buffered data per tag is lost. Add upload_on_shutdown (default false) and upload_on_shutdown_timeout (default 20s). On exit, every buffered file is sealed and uploaded from cb_exit, which runs on the pipeline thread after the output workers have been joined; the upstream stays in sync mode, so blocking requests are safe there. Per-request timeouts are clamped to the remaining budget, and files not uploaded in time stay in store_dir as before. Nothing is uploaded on hot reload, where the new instance recovers store_dir. Also create the upload timer in cb_worker_init. A backlog found at startup used to create the timer from cb_init on the engine scheduler, so every later upload ran on the pipeline thread, and the backlog was uploaded synchronously during init. With workers, the backlog is now uploaded from a one-shot timer on the worker shortly after start. Signed-off-by: amirgo1 <amir.gorodetzky@wiz.io>
Signed-off-by: amirgo1 <amir.gorodetzky@wiz.io>
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configuration
📒 Files selected for processing (3)
Included review availability: This review used your included allowance. Your plan provides up to 8 included reviews per hour; 6 remain after this review. 📝 WalkthroughWalkthroughThe GCS output plugin now supports configurable uploads during normal shutdown. It adds worker-aware upload queue initialization, retry handling, timeout validation, and tests for successful uploads, failures, timeouts, and multiple workers. ChangesGCS shutdown uploads
Priority: ⬇️ Low Estimated code review effort: 3 (Moderate) | ~20 minutes Change: Feature Sequence Diagram(s)sequenceDiagram
participant Exit as cb_gcs_exit
participant Store as store_dir
participant Queue as Upload queue
participant GCS
Exit->>Store: Queue and seal stored files
Exit->>Queue: Attempt uploads
Queue->>GCS: Send upload request
GCS-->>Queue: Return upload result
Queue->>Store: Delete successfully uploaded file
Merge Risk: 🔵 Low · up to Shutdown may take longer than the configured upload timeout when authentication or network I/O is slow. Clarify whether the limit is best-effort or enforce it across the full upload before merging. 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
out_gcsbuffers records instore_dirand uploads a file when it reachestotal_file_sizeorupload_timeout. On exit, whatever is still buffered stays instore_dirfor the next start. When the next start never comes on that disk, for example when a Kubernetes node is removed, up toupload_timeoutof data per tag is lost.This PR adds two options:
upload_on_shutdown(defaultfalse): upload every buffered file when Fluent Bit stops, instead of leaving it instore_dirfor the next start.upload_on_shutdown_timeout(default20s,0means no limit): the maximum time spent on those uploads. Files not uploaded in time stay instore_dir.The uploads run from
cb_exit, on the pipeline thread after the output workers have been joined, where the upstream is in sync mode and blocking requests are safe. Each request's connect and IO timeouts are clamped to the remaining budget. A file that fails to upload stays instore_dir, as today; withpreserve_data_orderingthe step stops at the first failure. Nothing is uploaded on hot reload, where the new instance recoversstore_dirright away. The upload logic ofprocess_upload_queue()is factored intoupload_queue_entry()and reused.This PR also moves the upload timer to
cb_worker_init. Whencb_initfound a backlog instore_dir, it created the timer on the engine scheduler and uploaded the backlog synchronously during init, so every later upload ran on the pipeline thread and a large backlog delayed startup. With workers, the timer is now created on the worker and the backlog is uploaded from a one-shot timer shortly after start. Without workers the behavior is unchanged.Enter
[N/A]in the box, if an item is not applicable to your change.Testing
Before we can approve your change; please submit the following in a comment:
If this is a change to packaging of containers or native binaries then please confirm it works for all targets.
ok-package-testlabel to test for all targets (requires maintainer to do).Documentation
fluent/fluent-bit-docs#2767
Backporting
Fluent Bit is licensed under Apache 2.0, by submitting this pull request I understand that this code will be released under the terms of that license.