使用Python將pdf文件轉換成word,csv

本文轉載自查看原文 2018-05-01 15:16 9338

一：下載所需要的庫

1 ：pdfminer 安裝庫命令 pip install pdfminer3k

pdfminer3k是pdfminer的Python 3端口。PDFMiner是從PDF文檔中提取信息的工具。與其他PDF相關的工具不同，它完全專注於獲取和分析文本數據。PDFMiner允許獲取頁面中文本的確切位置，以及其他信息，如字體或線條。它包含一個PDF轉換器，可以將PDF文件轉換為其他文本格式（如HTML）。它有一個可擴展的PDF解析器，可用於其他目的而不是文本分析。

2: docx 安裝庫命令 pip install python_docx

Python DocX目前是Python OpenXML的一部分，你可以用它打開Word 2007及以后的文檔，而用它保存的文檔可以在Microsoft Office 2007/2010, Microsoft Mac Office 2008, Google Docs, OpenOffice.org 3, and Apple iWork 08中打開。

from pdfminer.pdfparser import PDFParser, PDFDocument
from pdfminer.pdfinterp import PDFResourceManager, PDFPageInterpreter,process_pdf
from pdfminer.layout import LAParams
from pdfminer.converter import PDFPageAggregator
from pdfminer.pdfinterp import PDFTextExtractionNotAllowed
from docx import Document
document = Document()
import warnings
warnings.filterwarnings("ignore")
from pdfminer.converter import TextConverter
from pdfminer.layout import LAParams
from io import StringIO
from urllib.request import urlopen
import pandas as pd

def readPDF(pdfFile):
    rsrcmgr = PDFResourceManager()
    retstr = StringIO()
    laparams = LAParams()
    device = TextConverter(rsrcmgr, retstr, laparams=laparams)

    process_pdf(rsrcmgr, device, pdfFile)
    device.close()

    content = retstr.getvalue()
    retstr.close()
    return content
def save_to_file(file_name, contents):
    fh = open(file_name, 'w')
    fh.write(contents)
    fh.close()

save_to_file('mobiles.txt', 'your contents str')


def main():
    pdfFile = urlopen("http://pythonscraping.com/pages/warandpeace/chapter1.pdf")
    outputString = readPDF(pdfFile)
    #c.word
    save_to_file('c.csv',outputString)
if __name__ == '__main__':
    main()

使用docx 保存為word

from pdfminer.pdfparser import PDFParser, PDFDocument
from pdfminer.pdfinterp import PDFResourceManager, PDFPageInterpreter
from pdfminer.layout import LAParams
from pdfminer.converter import PDFPageAggregator
from pdfminer.pdfinterp import PDFTextExtractionNotAllowed
from docx import Document
document = Document()
import warnings
warnings.filterwarnings("ignore")
import os
file_name=os.open('/Users/dudu/Desktop/test1/a.pdf',os.O_RDWR )

def main():

    fn = open(file_name,'rb')
    parser = PDFParser(fn)
    doc = PDFDocument()
    parser.set_document(doc)
    doc.set_parser(parser)
    resource = PDFResourceManager()
    laparams = LAParams()
    device = PDFPageAggregator(resource,laparams=laparams)
    interpreter = PDFPageInterpreter(resource,device)
    for i in doc.get_pages():
        interpreter.process_page(i)
        layout = device.get_result()
        for out in layout:
            if hasattr(out,"get_text"):
                content = out.get_text().replace(u'\xa0', u' ') 
                document.add_paragraph(
                    content, style='ListBullet'   
                )
            document.save('a'+'.docx')
    print ('處理完成')
 
if __name__ == '__main__':
    main()

加下面的公眾號，我會定期發一些資料。

免責聲明！

本站轉載的文章為個人學習借鑒使用，本站對版權不負任何法律責任。如果侵犯了您的隱私權益，請聯系本站郵箱yoyou2525@163.com刪除。

猜您在找 python，多圖片轉換成pdf文件怎么將掃描版pdf文件怎么轉換成word文件 python 文件格式為 txt 轉換成 csv 格式 Json文件轉換成CSV 怎樣把txt文檔轉換成csv文件？ word轉換成pdf后，怎樣才能禁止復制和編輯呢？ vue文件流轉換成pdf預覽(pdf.js+iframe)&&使用vue-pdf實現pdf預覽 C#：CsvReader讀取.CSV文件（轉換成DataTable） dvi文件和將dvi文件轉換成pdf格式 DWG是什么文件？DWG文件如何轉換成PDF格式？