用户登录
用户注册

分享至

pythonword模板 python输出word内容

  • 作者: 巭孬嫑夯茓夯茓奣炛
  • 来源: 51数据库
  • 2020-04-21

程序导出word文档的方法

将web/html内容导出为world文档,再java中有很多解决方案,比如使用Jacob、Apache POI、Java2Word、iText等各种方式,以及使用freemarker这样的模板引擎这样的方式。php中也有一些相应的方法,但在python中将web/html内容生成world文档的方法是很少的。其中最不好解决的就是如何将使用js代码异步获取填充的数据,图片导出到word文档中。

1. unoconv

功能:

1.支持将本地html文档转换为docx格式的文档,所以需要先将网页中的html文件保存到本地,再调用unoconv进行转换。转换效果也不错,使用方法非常简单。

\# 安装

sudo apt-get install unoconv

\# 使用

unoconv -f pdf *.odt

unoconv -f doc *.odt

unoconv -f html *.odt

缺点:

1.只能对静态html进行转换,对于页面中有使用ajax异步获取数据的地方也不能转换(主要是要保证从web页面保存下来的html文件中有数据)。

2.只能对html进行转换,如果页面中有使用echarts,highcharts等js代码生成的图片,是无法将这些图片转换到word文档中;

3.生成的word文档内容格式不容易控制。

2. python-docx

功能:

1.python-docx是一个可以读写word文档的python库。

使用方法:

1.获取网页中的数据,使用python手动排版添加到word文档中。

from docx import Document

from docx.shared import Inches

document = Document()

document.add_heading('Document Title', 0)

p = document.add_paragraph('A plain paragraph having some ')

p.add_run('bold').bold = True

p.add_run(' and some ')

p.add_run('italic.').italic = True

document.add_heading('Heading, level 1', level=1)

document.add_paragraph('Intense quote', style='IntenseQuote')

document.add_paragraph(

'first item in unordered list', style='ListBullet'

)

document.add_paragraph(

'first item in ordered list', style='ListNumber'

)

document.add_picture('monty-truth.png', width=Inches(1.25))

table = document.add_table(rows=1, cols=3)

hdr_cells = table.rows[0].cells

hdr_cells[0].text = 'Qty'

hdr_cells[1].text = 'Id'

hdr_cells[2].text = 'Desc'

for item in recordset:

row_cells = table.add_row().cells

row_cells[0].text = str(item.qty)

row_cells[1].text = str(item.id)

row_cells[2].text = item.desc

document.add_page_break()

document.save('demo.docx')

from docx import Document

from docx.shared import Inches

document = Document()

for row in range(9):

t = document.add_table(rows=1,cols=1,style = 'Table Grid')

t.autofit = False #很重要!

w = float(row) / 2.0

t.columns[0].width = Inches(w)

document.save('table-step.docx')

缺点:

1.功能非常弱。有很多限制比如不支持模板等,只能生成简单格式的word文档。

程序导出PDF文档方法

1.pdfkit

功能:

1.wkhtmltopdf主要用于HTML生成PDF。

2.pdfkit是基于wkhtmltopdf的python封装,支持URL,本地文件,文本内容到PDF的转换,其最终还是调用wkhtmltopdf命令。是目前接触到的python生成pdf效果较好的。

优点:

1.wkhtmltopdf:利用webkit内核将HTML转为PDF

webkit是一个高效、开源的浏览器内核,包括Chrome和Safari在内的浏览器都使用了这个内核。Chrome打印当前网页的功能,其中有一个选项就是直接“保存为 PDF”。

2.wkhtmltopdf使用webkit内核的PDF渲染引擎来将HTML页面转换为PDF。高保真,转换质量很好,且使用非常简单。

使用方法:

\# 安装

pip install pdfkit

\# 使用

import pdfkit

pdfkit.from_url('', 'out.pdf')

pdfkit.from_file('test.html', 'out.pdf')

pdfkit.from_string('Hello!', 'out.pdf')

缺点:

1.对使用echarts,highcharts这样的js代码生成的图标无法转换为pdf(因为它的功能主要是将html转换为pdf,而不是将js转换为pdf)。对于纯静态页面的转换效果还是不错的。

2.其他

其他生成pdf的插件还有:weasyprint,reportlab,PyPDF2等,经简单试验都不如pdfkit效果好,且有些用法复杂。

如何用python或者R批量生成固定格式的word文档

office 2007中不能直接打开VB编辑器,请按Alt + F11打开。

import win32com.client # 导入脚本模块 WordApp = win32com.client.Dispatch("Word.Application") # 载入WORD模块

WordApp.Visible = True

# 显示Word应用程序

1、 新建Word文档

doc = WordApp.Documents.Add()

# 新建空文件

doc = WordApp.Documents.Open(r"d:\2011专业考试计划.doc") # 打开指定文档

doc.SaveAs(r"d:\2011专业考试计划.doc")

# 文档保存

doc.Close(-1)

# 保存后关闭,doc.Close()或doc.Close(0)直接关闭不保存

2、 页面设置

doc.PageSetup.PaperSize = 7

# 纸张大小, A3=6, A4=7

doc.PageSetup.PageWidth = 21*28.35 # 直接设置纸张大小, 使用该设置后PaperSize设置取消

doc.PageSetup.PageHeight = 29.7*28.35 # 直接设置纸张大小

doc.PageSetup.Orientation = 1 # 页面方向, 竖直=0, 水平=1 doc.PageSetup.TopMargin = 3*28.35

# 页边距上=3cm,1cm=28.35pt

doc.PageSetup.BottomMargin = 3*28.35 # 页边距下=3cm doc.PageSetup.LeftMargin = 2.5*28.35 # 页边距左=2.5cm doc.PageSetup.RightMargin = 2.5*28.35 # 页边距右=2.5cm

doc.PageSetup.TextColumns.SetCount(2) # 设置页面分栏=2

3、 格式设置

sel = WordApp.Selection

# 获取Selection对象 sel.InsertBreak(8)

# 插入分栏符=8, 分页符=7

sel.Font.Name = "黑体" # 字体 sel.Font.Size = 24 # 字大 sel.Font.Bold = True # 粗体 sel.Font.Italic = True # 斜体 sel.Font.Underline = True

# 下划线

sel.ParagraphFormat.LineSpacing = 2*12 # 设置行距,1行=12磅

sel.ParagraphFormat.Alignment = 1 # 段落对齐,0=左对齐,1=居中,2=右对齐 sel.TypeText("XXXX") # 插入文字 sel.TypeParagraph()

# 插入空行

注:ParagraphFormat属性必须使用TypeParagraph()之后才能二次生效

python操作word文档,如何合并单元格

>>>app=my.Office.Word.GetInstance()

>>>doc=app.Documents[0]

>>>table=doc.Tables[1]

>>>table.Cell(1,1).Select()

>>>app.Selection.MoveDown(Unit=5,Count=2,Extend=1)

>>>app.Selection.Cells.Merge()

>>>

  1. my.Office.Word.GetInstance()用win32com得到Word的Application对象的实例

  2. 我所使用的样本word文件中包含两个Table第二个Table是想要修改的

  3. table.Cell(1,1).Select()用于选中这个样表的第一个单元格

  4. app.Selection.MoveDown用于获得向下多选取3个单元格

  5. app.Selection.Cells.Merge()用于执行合并工作

python新建word文档

话说,你是在自己电脑上好好的,然后突然不行了

还是在别人电脑不行了?

word.displayalerts

这个是2013的属性

Microsoft Word 14.0,这是2010版更多

我一直在自己电脑上使用的,用的是2010版的office

电脑上同时安装了2013版的,但从来没用过啊。

哦,呵呵呵,这个我知道的,你装过2013,这个代码才跑得起来的

之前你word.application的时候,系统给你适配到了2013,你的程序就能跑

现在它不知道为啥给你适配到2010了,这个属性就没了

删掉这一行就是了

--我把这行删了,然后

AttributeError: '<win32com.gen_py.Microsoft Word 14.0 Object Library._Application instance at 0x52122704>' object has no attribute 'visible'

它连visible都识别不了了 飙泪ing。。。

额。。。那就不是这个问题,代码恢复回来

重启系统

刚试了,还是不行。。。

我刚查了一下,有人说这么解决:

问题解决方法:删除该库的.pyc文件,重新运行代码;或者找一个可以运行代码的环境,拷贝替换当前机器的.pyc文件即可

OK了!我将win32库重新安装了下,果然成功了!非常感谢!

怎么把python输出为word

程序导出word文档的方法

将web/html内容导出为world文档,再java中有很多解决方案,比如使用Jacob、Apache POI、Java2Word、iText等各种方式,以及使用freemarker这样的模板引擎这样的方式。php中也有一些相应的方法,但在python中将web/html内容生成world文档的方法是很少的。其中最不好解决的就是如何将使用js代码异步获取填充的数据,图片导出到word文档中。

1. unoconv

功能:

1.支持将本地html文档转换为docx格式的文档,所以需要先将网页中的html文件保存到本地,再调用unoconv进行转换。转换效果也不错,使用方法非常简单。

?

\# 安装

sudo apt-get install unoconv

\# 使用

unoconv -f pdf *.odt

unoconv -f doc *.odt

unoconv -f html *.odt

缺点:

1.只能对静态html进行转换,对于页面中有使用ajax异步获取数据的地方也不能转换(主要是要保证从web页面保存下来的html文件中有数据)。

2.只能对html进行转换,如果页面中有使用echarts,highcharts等js代码生成的图片,是无法将这些图片转换到word文档中;

3.生成的word文档内容格式不容易控制。

2. python-docx

功能:

1.python-docx是一个可以读写word文档的python库。

使用方法:

1.获取网页中的数据,使用python手动排版添加到word文档中。

如何用python读取word

使用Python的内部方法open()读取文本文件

try:

f=open('/file','r')

print(f.read())

finally:

iff:

f.close()

如果读取word文档推荐使用第三方插件,python-docx 可以在官网上下载

使用方式

#-*-coding:cp936-*-

importdocx

document=docx.Document(文件路径)

docText='\n\n'.join([

paragraph.text.encode('utf-8')forparagraphindocument.paragraphs

])

printdocText

python操作word文档表格

>>>app=my.Office.Word.GetInstance()

>>>doc=app.Documents[0]

>>>printdoc.Name

VBA工具集.doc

>>>doc.Tables.Count

2

>>>table=doc.Tables[1]

>>>table.Cell(1,1).Select()

>>>app.Selection.MoveEnd(Unit=12,Count=4)

4

>>>app.Selection.Cells.Shading.Texture=-10

>>>

1.my.Office.Word.GetInstance()用win32com得到Word的Application对象的实例

2.我所使用的样本word文件中包含两个Table第二个Table是想要修改的

3.table.Cell(1,1).Select()用于选中这个样表的第一个单元格

4.app.Selection.MoveEnd用于获得向右多选取4个单元格,wdCell=12,用于指示按单元格移动

5.app.Selection.Cells.Shading.Texture = -10用于执行阴影底纹的设置工作,wdTextureDiagonalUp=-10是一个代表斜向右上的底纹样式的常数

python word文件处理

#-*- encoding: utf8 -*-

import win32com

from win32com.client import Dispatch, constants

import win32com.client

import __main__

import os

import new

import sys

import re

import string

reload(sys)

sys.setdefaultencoding('utf8')

#from fileinput import filename

class Word(object):

#初始化word对象

def __init__(self, uri):

self.objectword(uri)

#创建word对象

def objectword(self,url):

self.word = win32com.client.Dispatch('Word.Application')

self.word.Visible = 0

self.word.DisplayAlerts = 0

self.docx = self.word.Documents.Open(url)

self.wrange = self.docx.Range(0, 0)

#关闭word

def close(self):

self.word.Documents.Close()

self.word.Quit()

#创建word

def create(self):

pass

#在word中进行查找

def findword(self, key):

question = []

uri = r'E:\XE\ctb.docx'

self.objectword(uri)

#读取所有的word文档内容

range = self.docx.Range(self.docx.Content.Start,self.docx.Content.End)

question = str(range).split("&")

#查找内容

#question = re.split(r"(\r[1][0-9][0-9]+.)",str(range))

#l = question[0].split("\d+.")

for questionLine in question:

questionLine = questionLine.strip('\n')

l = re.split(r"([1][0-9][0-9]+.)",questionLine)

del l[0]

for t in l:

s = str(key[0:3])

if str(t).find(s) > -1:

#插入

g = string.join(l)

print g.encode('gb2312')

#print g.decode("")

self.insertword(g)

print "sss"

else:

print "ttt"

#插入word

def insertword(self,w):

url = r'E:\XE\ctb.doc'

self.objectword(url)

self.wrange.InsertAfter(w)

pass

#读取数据源

def source(self, src):

f = open(src)

d = f.readlines()

for l in d:

name, question01, question02, question03, question04, question05 = tuple(l.decode('utf8').split('\t'))

if question01 != u'全对':

#self.wrange.InsertAfter(name)

self.findword(question01)

return self

Word(r'E:\XE\xx.docx').source(r'E:\XE\xe.txt').close()

转载请注明出处51数据库 » pythonword模板 python输出word内容

软件
前端设计
程序设计
Java相关