C#读取PDF文件的文本内容

本文转载自查看原文 2017-05-25 15:37 10165 C#（.Net）

  public static string ReadPdfContent(string filepath)
        {
            try
            {
                string pdffilename = filepath;
                PdfReader pdfReader = new PdfReader(pdffilename);
                int numberOfPages = pdfReader.NumberOfPages;
                StringBuilder text = new StringBuilder();
                for (int i = 1; i <= numberOfPages; ++i)
                {
                    text.Append(iTextSharp.text.pdf.parser.PdfTextExtractor.GetTextFromPage(pdfReader, i));
                }
                pdfReader.Close();
                return text.ToString();
            }
            catch (Exception ex)
            {
                return "原因：" + ex.ToString();
            }
        }

注：此方法需要引用iTextSharp

免责声明！

本站转载的文章为个人学习借鉴使用，本站对版权不负任何法律责任。如果侵犯了您的隐私权益，请联系本站邮箱yoyou2525@163.com删除。

猜您在找 C#读取PDF文件 C#读取PDF文档文字内容 C# 读取txt文本的内容 C# winfrom 读取txt文本内容 C# 读取文件内容 python 读取pdf文本内容 C#对文件下文本文件内容的读取 c# 打开文件和读取文件内容 C#异步将文本内容写入文件 C#读取大文本文件