Add outputEncoding parameter to enable encoding conversion (MRESOURCES-232) - #531
Conversation
gnodet-bot
left a comment
There was a problem hiding this comment.
Clean feature addition — well-structured parameter, good Javadoc, solid IT with both string and byte-level verification. One inaccuracy to fix on the @since tag.
The SNAPSHOT dependency on maven-filtering 4.0.0-beta-2-SNAPSHOT is acknowledged in the PR description and expected to be updated before merge.
This review was generated by an AI agent, Hermès on behalf of @gnodet.
a465a2f to
a6cef79
Compare
gnodet-bot
left a comment
There was a problem hiding this comment.
Re-review after force-push. The previous finding (@since 3.4.1 inaccuracy) is not addressed — the tag still references an unreleased version. Rest of the diff is clean.
This review was generated by an AI agent, Hermès on behalf of @gnodet.
a03a51a to
b3a5667
Compare
gnodet-bot
left a comment
There was a problem hiding this comment.
Re-review after rebase (a6cef79→b3a5667). The snapshot repository cleanup (.mvn/maven.config with -Puse-apache-snapshots instead of inline <repositories>) is a good improvement.
However, the @since tag finding raised in two previous reviews is still not addressed.
This review was generated by an AI agent, Hermès on behalf of @gnodet.
gnodet-bot
left a comment
There was a problem hiding this comment.
Re-review after c87ba70: the @since tag is now correctly set to 4.0.0-beta-2 — all three previous findings are addressed. Code is clean, Javadoc is accurate, IT covers both string content and raw byte verification.
This review was generated by an AI agent, Hermès on behalf of @gnodet.
gnodet-bot
left a comment
There was a problem hiding this comment.
Re-review after c87ba70. The @since tag finding raised in three previous reviews is now addressed — correctly updated from 3.4.1 to 4.0.0-beta-2.
Diff is clean: well-documented new parameter, proper null guard, solid IT with both string and byte-level encoding verification. No new issues.
This review was generated by an AI agent, Hermès on behalf of @gnodet.
rmannibucau
left a comment
There was a problem hiding this comment.
if we have a doc explaining the breaking text it can imply it can be neat to link in the javdoc - for common encodings not all
|
Thanks @rmannibucau — good suggestion. Added a paragraph listing common encodings and a |
gnodet-bot
left a comment
There was a problem hiding this comment.
Re-review after cb7c4fb: @rmannibucau's suggestion is addressed — new Javadoc paragraph lists common encodings (UTF-8, ISO-8859-1, US-ASCII, UTF-16) and links to java.nio.charset.Charset. Syntax is correct, HTML tags properly paired. No issues.
This review was generated by an AI agent, Hermès on behalf of @gnodet.
…SOURCES-232) - Add `outputEncoding` parameter to `ResourcesMojo` to configure the charset used for writing filtered resources, separate from the read charset (`encoding`) - Enables lossless encoding conversion (e.g. ISO-8859-1 input → UTF-8 output) without manually converting source files - Parameter also supports per-resource override via `<outputEncoding>` inside `<resource>` blocks (requires maven-filtering 4.0.0-beta-3+) - Only applies to filtered resources; non-filtered files are copied as raw bytes - Add IT `MRESOURCES-232` verifying Latin-1 source with filter token converts to UTF-8 Closes apache#335
cb7c4fb to
e1099a2
Compare
Summary
Adds a new
outputEncodingplugin parameter that configures the character encoding used to write filtered resources. When set, the plugin reads source files withencodingand writes filtered output withoutputEncoding, enabling lossless encoding conversion during the build.Closes #335
Companion filtering PR: apache/maven-filtering#407
Use case
Converting a legacy EBCDIC or Latin-1 (ISO-8859-1) source tree to UTF-8 output without manually converting the source files:
Or as a CLI property:
-Dmaven.resources.outputEncoding=UTF-8Per-resource encoding override via
<encoding>/<outputEncoding>inside a<resource>block is also supported (requires the companion filtering PR to be released first).Notes
<filtering>true</filtering>). Non-filtered files are always copied as raw bytes — encoding is irrelevant for them.outputEncodingis not set (the default), the same encoding is used for reading and writing, preserving full backward compatibility.pom.xmlreferencesmaven-filtering 4.0.0-beta-2-SNAPSHOTfor local build; will be updated to the released version once the companion PR is merged and released.Changes
ResourcesMojo: new@Parameter outputEncodingfield wired intoMavenResourcesExecution.setOutputEncoding()src/it/MRESOURCES-232/: IT with a Latin-1 source file containing a filter token and an é character; verifies the output file is UTF-8-encoded with the token expanded and é surviving as0xC3 0xA9Closes #335