Repository navigation
Conversation
|
Hi:
You could have a global constant dictionary of the conversion to and from for serialisation. I don't think it really matters.
I can't think of any now. I thing the raw counts of each failure will be the level I go to for inspection / research. But maybe with usage I'll be able to have more insights. |
adam-sutton-1992
left a comment
There was a problem hiding this comment.
LGTM - Perhaps test it on examples from NER and Linking one pipelines?
That lead me into a rabit hole to discover that this process kind of bricks the dataset aware implementation because it does postprocessing (i.e calls the pipe again) when determining some bits. So doing linker-only failure modes was just plain erroring out. So thanks for pointing that out! I'll be able to implement a fix for this. For reference, the "linker only failure modes" for this were: NOTE: the below threshold will also include a bunch for which the closest-similarity concept wasn't picked because it was outside the filter. EDIT: when I do "ner only failure modes" (for 1 doc) with Looks like that's because of the EDIT2: I think this is as good as we're going to be able to get with the NER-only use case. That's because of the nature of the NER-only step that assumes a perfect linker (that knows all the CUIs, and knows all their names). |
|
I've now made a few changes:
I've also reran the metrics and here's the results in
NOTE: The numbers have changed somewhat. This has to do with the fact that I now used the real pipe's output documents/entities/tokens rather than rerunning over context. |
This PR (optionally) adds failure modes to false negatives in the output of the
get_statsmethod.The list of failure modes used is in the source, but I'll bring it up here as well for clarity:
List of failure modes and their brief descriptions
Copied from code as of 2026-10-06 at 15-59.
The idea would be that if the model fails to corectly identify a specific concept, this should help identify why the specific issues are.
For reference, with the following code:
and the 2025 model running against the Snomed CT Entity Linking challenge dataset, we get the output of:
There are a number of options, some of which allow a lot more informaton to be shown in terms of the examples.
I will give a short example of more detailed output for 1 document from the same dataset
The code used was almost the same:
And the output we get (only in terms of summarisation) is:
A few caveats / questions / decisions:
Enumin the final output so as to allow easier serialisation?