Repository navigation
fix: preserve a leading BOM when set_key and unset_key rewrite a file - #687
MohammedAlkindi wants to merge 4 commits into
Conversation
The parser strips a leading UTF-8 BOM so the first variable is read correctly, but the stripped BOM never reaches the bindings that set_key and unset_key write back out. Rewriting a file that had one therefore dropped it silently, changing a part of the file the caller did not ask to touch. Carry it across in rewrite(), which both writers share.
Windows validation of PR #687I used an AI-assisted local test script on Windows with Python 3.12.10 to The matrix covered 8 operations for each of 4 encoding configurations: All 32 parsed-value checks passed on each revision. For a UTF-8-BOM file One existing edge case remains: with import codecs
from pathlib import Path
from tempfile import TemporaryDirectory
from dotenv import unset_key
with TemporaryDirectory() as folder:
path = Path(folder) / "sample.env"
path.write_text("FIRST=value\n", encoding="utf-8-sig")
assert path.read_bytes().startswith(codecs.BOM_UTF8)
unset_key(path, "FIRST", encoding="utf-8-sig")
print(repr(path.read_bytes())) # b'' on both revisionsThis was a focused local matrix, not a full-suite or CI run. I have not |
…-rewrite # Conflicts: # CHANGELOG.md
…-rewrite # Conflicts: # CHANGELOG.md
set_keyandunset_keydrop a leading UTF-8 BOM from files that have one. #640 made the parser strip the BOM so the first variable parses, but the stripped BOM never reaches the bindings, so the rewrite writes the file back without it:rewrite()now carries the BOM across, so both writers keep it, and only files that already had one change behaviour. The check is encoding-safe: a latin-1 file never decodes to\ufeff, and utf-16 handles its own BOM in the codec.The new
set_keyandunset_keytests fail on main on the missing BOM and pass with the fix. The full suite passes, andruff checkandruff format --checkare clean.Related: #640, #637