Bug report
Bug description:
The C implementation of time.fromisoformat() and datetime.fromisoformat() accepts stray text between the time and a UTC offset or Z. The pure-Python implementation rejects the same strings:
>>> from datetime import time, datetime
>>> time.fromisoformat('12:30:45 +02:00')
datetime.time(12, 30, 45, tzinfo=datetime.timezone(datetime.timedelta(seconds=7200)))
>>> time.fromisoformat('12:30x+05:00')
datetime.time(12, 30, tzinfo=datetime.timezone(datetime.timedelta(seconds=18000)))
>>> time.fromisoformat('1230:Z')
datetime.time(12, 30, tzinfo=datetime.timezone.utc)
>>> time.fromisoformat('12:30:45.123456 junk Z')
datetime.time(12, 30, 45, 123456, tzinfo=datetime.timezone.utc)
>>> datetime.fromisoformat('2020-01-01T12:30x+05:00')
datetime.datetime(2020, 1, 1, 12, 30, tzinfo=datetime.timezone(datetime.timedelta(seconds=18000)))
>>> import _pydatetime # Lib/_pydatetime.py from main
>>> _pydatetime.time.fromisoformat('12:30:45 +02:00')
ValueError: Invalid isoformat string: '12:30:45 +02:00'
parse_hh_mm_ss_ff() in Modules/_datetimemodule.c returns 1 ("not at end") in two cases:
- one character remains before
tstr_end after HH, MM or SS;
- any text follows a fraction of 6 or more digits.
It also returns 1 for every valid string with an offset, because it reads the designator itself as c. So parse_isoformat_time() can only treat rv == 1 as an error when there is no offset. When an offset follows, the stray text is accepted silently.
A comment on gh-130959 reported the '...05.600000 +02:30' form, but that issue was closed after the pure-Python fix, and the C side is unchanged. gh-107779 ('20230808120000Z' losing a digit) goes through the same rv == 1 path.
Checked on 3.11.15 with the C accelerator (output above), and _pydatetime from main (7eada7c). The C code on main was read, not built; the code path is unchanged.
Suggested fix: have parse_hh_mm_ss_ff() return 0 when c is the terminator or the designator (p > p_end) and 1 when it is a stray character (p == p_end), and return p != p_end after the fraction. parse_isoformat_time() can then reject rv == 1 whether or not an offset follows. I have a draft patch with tests. It is not compiled yet; I checked it through a Python transliteration of the C code against the fromisoformat vectors in datetimetester.py, and every example there still parses. I can open a PR if this is accepted.
CPython versions tested on:
3.11, CPython main branch
Operating systems tested on:
Windows
Bug report
Bug description:
The C implementation of
time.fromisoformat()anddatetime.fromisoformat()accepts stray text between the time and a UTC offset orZ. The pure-Python implementation rejects the same strings:parse_hh_mm_ss_ff()inModules/_datetimemodule.creturns 1 ("not at end") in two cases:tstr_endafterHH,MMorSS;It also returns 1 for every valid string with an offset, because it reads the designator itself as
c. Soparse_isoformat_time()can only treatrv == 1as an error when there is no offset. When an offset follows, the stray text is accepted silently.A comment on gh-130959 reported the
'...05.600000 +02:30'form, but that issue was closed after the pure-Python fix, and the C side is unchanged. gh-107779 ('20230808120000Z'losing a digit) goes through the samerv == 1path.Checked on 3.11.15 with the C accelerator (output above), and
_pydatetimefrommain(7eada7c). The C code onmainwas read, not built; the code path is unchanged.Suggested fix: have
parse_hh_mm_ss_ff()return 0 whencis the terminator or the designator (p > p_end) and 1 when it is a stray character (p == p_end), and returnp != p_endafter the fraction.parse_isoformat_time()can then rejectrv == 1whether or not an offset follows. I have a draft patch with tests. It is not compiled yet; I checked it through a Python transliteration of the C code against thefromisoformatvectors indatetimetester.py, and every example there still parses. I can open a PR if this is accepted.CPython versions tested on:
3.11, CPython main branch
Operating systems tested on:
Windows