文章详情页

文本处理 - 求教使用python库提取pdf的方法？

浏览：300日期：2022-09-02 11:18:15

问题描述

使用过pypdf 对英文pdf文档处理比较简单，但是对中文的支持好像不太好

使用过textract 看文档支持的格式比较多方法也比较简单，但是老师出错

-- coding: utf-8 --

import textractimport pyPdfimport pdf2textimport pdfminerimport chardet

text = textract.process('F:ll.pdf',method = ’pdfminer’)print text

这个出错是编码问题-- coding: utf-8 --

import textractimport pyPdfimport pdfminerimport chardet

text = textract.process('F:ll.pdf',method = ’pdfminer’)print text

这个出错类型不清楚

少使用了pdf2text库，但是出错情况好像不一样。

pdfminer库还没看过，看着好像麻烦一些，求解一下解析提取中文的pdf的方法。谢谢

问题解答

回答1：

之前用过的pdfminer pip install pdfminer

# -*- coding: utf-8 -*-from bs4 import BeautifulSoupimport requestsimport refrom pdfminer.pdfinterp import PDFResourceManager, PDFPageInterpreterfrom pdfminer.converter import TextConverterfrom pdfminer.layout import LAParamsfrom cStringIO import StringIO#from io import StringIO for python3from io import openfrom pdfminer.pdfpage import PDFPagedef pdf_txt(url): rsrcmgr = PDFResourceManager() retstr = StringIO() codec = ’utf-8’ laparams = LAParams() device = TextConverter(rsrcmgr, retstr, codec=codec, laparams=laparams) f = requests.get(url).content fp = StringIO(f) interpreter = PDFPageInterpreter(rsrcmgr, device) password = '' maxpages = 0 caching = True pagenos = set() for page in PDFPage.get_pages(fp, pagenos, maxpages=maxpages, password=password, caching=caching, check_extractable=True):interpreter.process_page(page) fp.close() device.close() str = retstr.getvalue() retstr.close() return strtxt=tpdf_txt(’http://pythonscraping.com/pages/warandpeace/chapter1.pdf’)print txt#如果pdf含有中文，输出到文件#open(’pdf.txt’,’wb’).write(txt)python readpdf.py’’’CHAPTER I'Well, Prince, so Genoa and Lucca are now just family estates oftheBuonapartes. But I warn you, if you don’t tell me that thismeans war,if you still try to defend the infamies and horrorsperpetrated bythat Antichrist- I really believe he is Antichrist- I willhavenothing more to do with you and you are no longer my friend,no longermy ’faithful slave,’ as you call yourself! But how do youdo? I seeI have frightened you- sit down and tell me all the news.'It was in July, 1805, and the speaker was the well-knownAnnaPavlovna Scherer, maid of honor and favorite of theEmpress MaryaFedorovna. With these words she greeted PrinceVasili Kuragin, a manof high rank and importance, who was thefirst to arrive at herreception. Anna Pavlovna had had a cough forsome days. She was, asshe said, suffering from la grippe; grippebeing then a new word inSt. Petersburg, used only by the elite.All her invitations without exception, written in French,anddelivered by a scarlet-liveried footman that morning, ran as’’’

Python 编程

上一条：flask - python jinja2 如何从获取javascript function(_index) 传过来的参数 index?下一条：python - for循环print怎样才能输出csv呢

相关文章：

1. javascript - 如何将psd裁切下来的图片非常清晰的宣示出来2. python小白，关于函数问题3. docker gitlab 如何git clone？4. angular.js - angular中的a标签不起作用5. javascript - weex和node,js到底是怎样一个关系呢？6. java - ssm框架jar包如何区分和管理？7. docker内创建jenkins访问另一个容器下的服务器问题8. 前端 - 类到底该如何去命名 .newsList 这种的命名难道真的不是过度语义化吗？~9. javascript - 关于js原生事件的绑定与解除绑定10. docker绑定了nginx端口外部访问不到

排行榜

					
					angular.js - angular中的a标签不起作用
javascript - weex和node,js到底是怎样一个关系呢？
docker镜像push报错
javascript - 关于js原生事件的绑定与解除绑定
dockerfile - 为什么docker容器启动不了？
docker gitlab 如何git clone？
java - ssm框架jar包如何区分和管理？
docker内创建jenkins访问另一个容器下的服务器问题
javascript - 如何将psd裁切下来的图片非常清晰的宣示出来
前端 - 类到底该如何去命名 .newsList 这种的命名难道真的不是过度语义化吗？~
python小白，关于函数问题
				

热门标签