In this repo, you can find Python Notebooks and sample model outputs we created for the information extraction use case in the journalism benchmark cookbook.
Info_Extraction_Scenario-1.ipynb demonstrates how we evaluated models' capability to extract unstructured information from pdf files for internal usage with FOIA data.
Info_Extraction_Scenario-2.ipynb demonstrates how we evaluated models' capability to extract unstructured information from plain text for internal usage with app store review data.
Info_Extraction_Scenario-3.ipynb demonstrates how we evaluated models' capability to extract structured information from images for internal usage with restricted items on Amazon data.
You can also find a sample output we collected from models in sample_model_outputs
If you have any questions, you can contact Charlotte Li