Repository navigation
Markdown Feeds: Decode HTML entities in titles, tagline and excerpts - #1086
Conversation
Post titles, the site tagline and feed excerpts were written into the Markdown document with their HTML entities intact (It’s, Tips & Tricks), while the converted post content and the site name next to them were already plain text. Titles were passed through wp_specialchars_decode(), which only reverses the five special characters, and the tagline and excerpts were not decoded at all. Decode all of them with one helper on the converter so the whole document is entity-free.
|
The following accounts have interacted with this PR and/or linked issues. I will continue to update these lists as activity occurs. You can also manually ask me to refresh this list by adding the If you're merging code through a pull request on GitHub, copy and paste the following into the bottom of the merge commit message. To understand the WordPress project's expectations around crediting contributors, please review the Contributor Attribution page in the Core Handbook. |
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## develop #1086 +/- ##
=============================================
+ Coverage 81.53% 81.55% +0.01%
- Complexity 3071 3072 +1
=============================================
Files 129 129
Lines 12253 12255 +2
=============================================
+ Hits 9991 9994 +3
+ Misses 2262 2261 -1
Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
|
@dkotter this might be worth squeezing into 1.4.0 |
Use get_bloginfo( 'charset' ) instead of a hardcoded UTF-8 when decoding entities in titles, the tagline and excerpts, so the decoded text is in the charset the Markdown response declares. Nothing changes on UTF-8 sites.
Test ReportPatch tested: #1086 (commit 556223c), compared with Environment
Steps
Results
The converted post content ( ✅ The patch fixes #1085. |


What?
Closes #1085
Decodes HTML entities in the plain-text parts of the Markdown Feeds output: post titles, the site tagline and feed excerpts.
Why?
On
developthose parts are served with entities left in, while the converted content and the site name next to them are plain text:Titles go through
wp_specialchars_decode( ..., ENT_QUOTES ), which only reverses the five special characters, so everythingwptexturizeproduces (’,“,–,…) stays encoded. The tagline and excerpts are not decoded at all. Any title with an apostrophe or quotes is affected.How?
Markdown_Converter::decode_entities(), a thin wrapper aroundhtml_entity_decode()withENT_QUOTES | ENT_HTML5and the site's charset.wp_strip_all_tags(), so an entity-encoded tag cannot turn into a real one and take text with it.Use of AI Tools
AI assistance: Yes
Tool(s): Claude Code and Codex
Used for: Investigation, implementation, tests, and PR wording. I reviewed the reasoning and test results, and I take responsibility for the contribution.
Testing Instructions
Tips & Tricksand publish a post titledIt's a "quoted" title./feed/markdown/and the post URL with?output_format=markdown. The tagline readsTips & Tricksand the title readsIt’s a “quoted” title, with the entities decoded.Automated:
npm run test:php -- --filter Markdown_passes with 35 tests and 77 assertions. Five tests are new. The three renderer tests fail ondevelopwith the entity strings above and pass here.WP_MULTISITE=1.composer lintand PHPStan pass.Also verified by hand on a local WordPress 7.1 site through
/feed/markdown/and?output_format=markdown, before and after.Changelog Entry