Bug report
Bug description:
date.fromisoformat() accepts a 10-character string in basic format and silently drops its last two characters:
>>> from datetime import date, datetime
>>> date.fromisoformat('2020010112') # e.g. a YYYYMMDDHH stamp
datetime.date(2020, 1, 1)
>>> date.fromisoformat('2020W011xx')
datetime.date(2019, 12, 30)
>>> datetime.fromisoformat('2020010112') # the datetime parser rejects it
ValueError: Invalid isoformat string: '2020010112'
date.fromisoformat admits inputs of length 7, 8 or 10. Both parse_isoformat_date() in Modules/_datetimemodule.c and _parse_isoformat_date() in Lib/_pydatetime.py read fixed-width fields from the start of the string, and neither checks that the whole string was consumed. A 10-character string without - at index 4 is therefore parsed as the 8-character basic date YYYYMMDD (or YYYYWwwD), and the trailing characters are ignored.
Such a string is not an ISO 8601 date, and it is not one of the documented exceptions. Because the result is a plausible date, a truncated or mistyped input (such as a 10-digit YYYYMMDDHH) goes unnoticed.
Checked on:
- 3.11.15 with the C accelerator: the output above;
Lib/_pydatetime.py from main (7eada7c), run under 3.11.15: same results for both strings;
Modules/_datetimemodule.c on main: read, not built. The code path is unchanged.
Suggested fix: after the last field, require that the whole string has been consumed (p - dtstr == len in C, len(dtstr) == pos in Python). I have a draft patch with tests for both implementations and can open a PR if this is accepted.
Found by differential testing fromisoformat against an independent ISO 8601 implementation. I found no existing report; the nearest is gh-152204, which covers other pure-Python-only inputs.
CPython versions tested on:
3.11, CPython main branch
Operating systems tested on:
Windows
Bug report
Bug description:
date.fromisoformat()accepts a 10-character string in basic format and silently drops its last two characters:date.fromisoformatadmits inputs of length 7, 8 or 10. Bothparse_isoformat_date()inModules/_datetimemodule.cand_parse_isoformat_date()inLib/_pydatetime.pyread fixed-width fields from the start of the string, and neither checks that the whole string was consumed. A 10-character string without-at index 4 is therefore parsed as the 8-character basic dateYYYYMMDD(orYYYYWwwD), and the trailing characters are ignored.Such a string is not an ISO 8601 date, and it is not one of the documented exceptions. Because the result is a plausible date, a truncated or mistyped input (such as a 10-digit
YYYYMMDDHH) goes unnoticed.Checked on:
Lib/_pydatetime.pyfrommain(7eada7c), run under 3.11.15: same results for both strings;Modules/_datetimemodule.conmain: read, not built. The code path is unchanged.Suggested fix: after the last field, require that the whole string has been consumed (
p - dtstr == lenin C,len(dtstr) == posin Python). I have a draft patch with tests for both implementations and can open a PR if this is accepted.Found by differential testing
fromisoformatagainst an independent ISO 8601 implementation. I found no existing report; the nearest is gh-152204, which covers other pure-Python-only inputs.CPython versions tested on:
3.11, CPython main branch
Operating systems tested on:
Windows