C# 使用Aspose.Pdf读取Pdf表格

本文转载自查看原文 2021-05-11 13:31 3449 技术

pdf文件内容：

1.上面我们有一个pdf文件，内容为表格形式的，在使用Aspose.Pdf读取的时候，如果不定义读取时的TextExtractionOptions，我们看一下读取的内容是什么样子的?

可以看到在读取pdf文字的时候，并没有按照表格划分，而是视觉上同一行的文字被划分到同一行，这样在处理数据的时候就比较麻烦。

2.我们定义一下TextExtractionOptions试试：

var textExtractionOptions = new TextExtractionOptions(TextExtractionOptions.TextFormattingMode.Raw);

我们可以看到这时候的文字已经按照单元格分开了。

参考代码：

 1 using Aspose.Pdf;
 2 using Aspose.Pdf.Text;
 3 using Aspose.Pdf.Text.TextOptions;
 4 
 5 namespace Test
 6 {
 7     class Program
 8     {
 9         static void Main(string[] args)
10         {
11             Document pdfDocument = new Document(@"d:\pdf.pdf");
12             var textExtractionOptions = new TextExtractionOptions(TextExtractionOptions.TextFormattingMode.Raw);
13             var textSearchOptions = new TextSearchOptions(true);
14             TextAbsorber textAbsorber = new TextAbsorber(textExtractionOptions, textSearchOptions);
15             pdfDocument.Pages.Accept(textAbsorber);
16             string content = textAbsorber.Text;
17         }
18     }
19 }

免责声明！

本站转载的文章为个人学习借鉴使用，本站对版权不负任何法律责任。如果侵犯了您的隐私权益，请联系本站邮箱yoyou2525@163.com删除。

猜您在找 Aspose.Pdf合并PDF文件 C# -- 使用Aspose.Cells创建和读取Excel文件 Aspose.Cells 插件使用 For C# c#使用Aspose打印文件 C# 绘制PDF嵌套表格 C#读取Excel表格的数据 C# 利用Aspose.Words .dll将本地word文档转化成pdf C#使用Aspose.Words把 word转成图片 C#+Aspose转office文档为PDF 使用aspose words把word文件转为pdf