Skip to content

Add outputEncoding parameter to enable encoding conversion (MRESOURCES-232) - #531

Merged
gnodet merged 1 commit into
apache:masterfrom
gnodet:feat/output-encoding
Sep 30, 2026
Merged

gnodet merged 1 commit into
apache:masterfrom
gnodet:feat/output-encoding

Conversation

@gnodet

@gnodet gnodet commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Adds a new outputEncoding plugin parameter that configures the character encoding used to write filtered resources. When set, the plugin reads source files with encoding and writes filtered output with outputEncoding, enabling lossless encoding conversion during the build.

Closes #335

Companion filtering PR: apache/maven-filtering#407

Use case

Converting a legacy EBCDIC or Latin-1 (ISO-8859-1) source tree to UTF-8 output without manually converting the source files:

<plugin>
  <groupId>org.apache.maven.plugins</groupId>
  <artifactId>maven-resources-plugin</artifactId>
  <configuration>
    <encoding>ISO-8859-1</encoding>
    <outputEncoding>UTF-8</outputEncoding>
  </configuration>
</plugin>

Or as a CLI property: -Dmaven.resources.outputEncoding=UTF-8

Per-resource encoding override via <encoding> / <outputEncoding> inside a <resource> block is also supported (requires the companion filtering PR to be released first).

Notes

  • The parameter only applies to filtered resources (<filtering>true</filtering>). Non-filtered files are always copied as raw bytes — encoding is irrelevant for them.
  • When outputEncoding is not set (the default), the same encoding is used for reading and writing, preserving full backward compatibility.
  • pom.xml references maven-filtering 4.0.0-beta-2-SNAPSHOT for local build; will be updated to the released version once the companion PR is merged and released.

Changes

  • ResourcesMojo: new @Parameter outputEncoding field wired into MavenResourcesExecution.setOutputEncoding()
  • src/it/MRESOURCES-232/: IT with a Latin-1 source file containing a filter token and an é character; verifies the output file is UTF-8-encoded with the token expanded and é surviving as 0xC3 0xA9

Closes #335

@gnodet-bot gnodet-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Clean feature addition — well-structured parameter, good Javadoc, solid IT with both string and byte-level verification. One inaccuracy to fix on the @since tag.

The SNAPSHOT dependency on maven-filtering 4.0.0-beta-2-SNAPSHOT is acknowledged in the PR description and expected to be updated before merge.

This review was generated by an AI agent, Hermès on behalf of @gnodet.

Comment thread src/main/java/org/apache/maven/plugins/resources/ResourcesMojo.java Outdated

@gnodet-bot gnodet-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-review after force-push. The previous finding (@since 3.4.1 inaccuracy) is not addressed — the tag still references an unreleased version. Rest of the diff is clean.

This review was generated by an AI agent, Hermès on behalf of @gnodet.

Comment thread src/main/java/org/apache/maven/plugins/resources/ResourcesMojo.java Outdated
@gnodet
gnodet force-pushed the feat/output-encoding branch from a03a51a to b3a5667 Compare September 27, 2026 23:47

@gnodet-bot gnodet-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-review after rebase (a6cef79→b3a5667). The snapshot repository cleanup (.mvn/maven.config with -Puse-apache-snapshots instead of inline <repositories>) is a good improvement.

However, the @since tag finding raised in two previous reviews is still not addressed.

This review was generated by an AI agent, Hermès on behalf of @gnodet.

Comment thread src/main/java/org/apache/maven/plugins/resources/ResourcesMojo.java Outdated

@gnodet-bot gnodet-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-review after c87ba70: the @since tag is now correctly set to 4.0.0-beta-2 — all three previous findings are addressed. Code is clean, Javadoc is accurate, IT covers both string content and raw byte verification.

This review was generated by an AI agent, Hermès on behalf of @gnodet.

@gnodet-bot gnodet-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-review after c87ba70. The @since tag finding raised in three previous reviews is now addressed — correctly updated from 3.4.1 to 4.0.0-beta-2.

Diff is clean: well-documented new parameter, proper null guard, solid IT with both string and byte-level encoding verification. No new issues.

This review was generated by an AI agent, Hermès on behalf of @gnodet.

@rmannibucau rmannibucau left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if we have a doc explaining the breaking text it can imply it can be neat to link in the javdoc - for common encodings not all

@gnodet

gnodet commented Sep 30, 2026

Copy link
Copy Markdown
Contributor Author

Thanks @rmannibucau — good suggestion. Added a paragraph listing common encodings and a {@link java.nio.charset.Charset} reference in cb7c4fb.

@gnodet-bot gnodet-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-review after cb7c4fb: @rmannibucau's suggestion is addressed — new Javadoc paragraph lists common encodings (UTF-8, ISO-8859-1, US-ASCII, UTF-16) and links to java.nio.charset.Charset. Syntax is correct, HTML tags properly paired. No issues.

This review was generated by an AI agent, Hermès on behalf of @gnodet.

@gnodet gnodet added enhancement New feature or request java Pull requests that update Java code labels Sep 30, 2026
…SOURCES-232)

- Add `outputEncoding` parameter to `ResourcesMojo` to configure the charset used
  for writing filtered resources, separate from the read charset (`encoding`)
- Enables lossless encoding conversion (e.g. ISO-8859-1 input → UTF-8 output)
  without manually converting source files
- Parameter also supports per-resource override via `<outputEncoding>` inside
  `<resource>` blocks (requires maven-filtering 4.0.0-beta-3+)
- Only applies to filtered resources; non-filtered files are copied as raw bytes
- Add IT `MRESOURCES-232` verifying Latin-1 source with filter token converts to UTF-8

Closes apache#335
@gnodet
gnodet force-pushed the feat/output-encoding branch from cb7c4fb to e1099a2 Compare September 30, 2026 07:16
@gnodet
gnodet merged commit d6f8666 into apache:master Sep 30, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request java Pull requests that update Java code

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[MRESOURCES-232] Resource copy filtering should use different encoding for source and output

3 participants