Skip to content

Problem indexing spreadsheets containing cells spanning multiple rows or columns #2159

Description

@stuzart

As the following screenshots show, when being converted to PDF (before being converted to text), the converter doesn't recognise the cell boundries, so just reads the text horizontally. This leads to incorrect text being indexed for search
There are similar issues where a cell may only be one row in height, but bleeds into the ajoining cell to the right.

e.g.
Cells that look like:

Image

get converted to the following text:

Image

There is a (possible) fix to the gem used documentcloud/docsplit#132 but was never released as a new gem. The are other potential alternatives, but some have issues registered describing a similar problem:

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

Status
No Status

Relationships

None yet

Development

No branches or pull requests

Issue actions