Skip to content

Releases: Unstructured-IO/unstructured

0.16.23

20 Feb 13:31
0df50fe
Compare
Choose a tag to compare

0.16.23

Enhancements

Features

Fixes

  • Fixes detect_filetype when SpooledTemporaryFile is passed. Previously some random name would get assigned to the file and the function raised error.

0.16.22

20 Feb 01:11
147add9
Compare
Choose a tag to compare

0.16.22

Enhancements

Features

Fixes

  • Fix open CVES in and bump dependencies

0.16.21

17 Feb 16:01
3403db1
Compare
Choose a tag to compare

Enhancements

  • Use password to load PDF with all modes

  • use vectorized logic to merge inferred and extracted layouts. Using the new LayoutElements data structure and numpy library to refactor the layout merging logic to improve compute performance as well as making logic more clear

  • Add PDF Miner configuration Now PDF Miner can be configured via pdfminer_line_overlap, pdfminer_word_margin, pdfminer_line_margin and pdfminer_char_margin parameters added to partition method.

Features

Fixes

  • Fix file type detection for NDJSON files NDJSON files were being detected as JSON due to having the same mime-type.

0.16.20

06 Feb 06:12
b10379c
Compare
Choose a tag to compare

0.16.20

Enhancements

Features

Fixes

  • Fix a security issue where rst and org files could read files in the local filesystem. Certain filetypes could 'include' or 'import' local files into their content, allowing partitioning of arbitrary files from the local filesystem. Partitioning of these files is now sandboxed.

0.16.19

05 Feb 17:21
5852260
Compare
Choose a tag to compare

Enhancements

Features

Fixes

  • Fix a bug where table extraction is skipped when it shouldn't. Pages with just one table as its content or starts with a table misses table extraction. The routing logic is now fixed.
  • Correct deprecated ruff invocation in make tidy. This will future-proof it or avoid surprises if someone happens to upgrade Ruff.
  • Remove upper bound constraint on python version in setup.py. Python3.13 is not yet officially supported, but allow users to try.
  • Fixes removing HTML elements from the inside of table cells in html partition v=2.0. The HTML partitioner now correctly preserves HTML elements from the inside of table cells.

0.16.17

29 Jan 12:52
55debaf
Compare
Choose a tag to compare

0.16.17

Enhancements

  • Refactoring the VoyageAI integration to use voyageai package directly, allowing extra features.

Features

Fixes

  • Fix a bug where build_layout_elements_from_cor_regions incorrectly joins texts in wrong order.

Full Changelog: 0.16.16...0.16.17

0.16.16

27 Jan 23:30
a447b81
Compare
Choose a tag to compare

0.16.16

Enhancements

Features

  • Vectorize layout (inferred, extracted, and OCR) data structure Using np.ndarray to store a group of layout elements or text regions instead of using a list of objects. This improves the memory efficiency and compute speed around layout merging and deduplication.

Fixes

  • Add auto-download for NLTK for Python Enviroment When user import tokenize, It will automatic download nltk data from tokenize.py file. Added AUTO_DOWNLOAD_NLTK flag in tokenize.py to download NLTK_DATA.
  • Correctly patch pdfminer to avoid PDF repair. The patch applied to pdfminer's parser caused it to occasionally split tokens in content streams, throwing PDFSyntaxError. Repairing these PDFs sometimes failed (since they were not actually invalid) resulting in unnecessary OCR fallback.
  • Drop usage of ndjson dependency

0.16.15

23 Jan 04:51
8d0b68a
Compare
Choose a tag to compare
  • Update unstructured-inference to 0.8.6 in requirements which removed layoutparser dependency libs
  • Update pdfminer-six to 20240706

0.16.14

20 Jan 13:00
efd9f64
Compare
Choose a tag to compare

Enhancements

Features

Fixes

  • Fix an issue with multiple values for infer_table_structure when paritioning email with image attachements the kwarg calls into partition to partition the image already contains infer_table_structure. Now partition function checks if the kwarg has infer_table_structure already

0.16.13

13 Jan 15:40
38eb661
Compare
Choose a tag to compare

Enhancements

  • Add character-level filtering for tesseract output. It is controllable via TESSERACT_CHARACTER_CONFIDENCE_THRESHOLD environment variable.

Features

Fixes

  • Fix NLTK Download to use nltk assets in docker image
  • removed the ability to automatically download nltk package if missing