python下載微信公眾號(hào)相關(guān)文章
本文實(shí)例為大家分享了python下載微信公眾號(hào)相關(guān)文章的具體代碼,供大家參考,具體內(nèi)容如下
目的:從零開始學(xué)自動(dòng)化測(cè)試公眾號(hào)中下載“pytest"一系列文檔
1、搜索微信號(hào)文章關(guān)鍵字搜索
2、對(duì)搜索結(jié)果前N頁(yè)進(jìn)行解析,獲取文章標(biāo)題和對(duì)應(yīng)URL
主要使用的是requests和bs4中的Beautifulsoup
Weixin.py
import requests from urllib.parse import quote from bs4 import BeautifulSoup import re from WeixinSpider.HTML2doc import MyHTMLParser class WeixinSpider(object): def __init__(self, gzh_name, pageno,keyword): self.GZH_Name = gzh_name self.pageno = pageno self.keyword = keyword.lower() self.page_url = [] self.article_list = [] self.headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/70.0.3538.110 Safari/537.36'} self.timeout = 5 # [...] 用來(lái)表示一組字符,單獨(dú)列出:[amk] 匹配 'a','m'或'k' # re+ 匹配1個(gè)或多個(gè)的表達(dá)式。 self.pattern = r'[\\/:*?"<>|\r\n]+' def get_page_url(self): for i in range(1,self.pageno+1): # https://weixin.sogou.com/weixin?query=從零開始學(xué)自動(dòng)化測(cè)試&_sug_type_=&s_from=input&_sug_=n&type=2&page=2&ie=utf8 url = "https://weixin.sogou.com/weixin?query=%s&_sug_type_=&s_from=input&_sug_=n&type=2&page=%s&ie=utf8" \ % (quote(self.GZH_Name),i) self.page_url.append(url) def get_article_url(self): article = {} for url in self.page_url: response = requests.get(url,headers=self.headers,timeout=self.timeout) result = BeautifulSoup(response.text, 'html.parser') articles = result.select('ul[class="news-list"] > li > div[class="txt-box"] > h3 > a ') for a in articles: # print(a.text) # print(a["href"]) if self.keyword in a.text.lower(): new_text=re.sub(self.pattern,"",a.text) article[new_text] = a["href"] self.article_list.append(article) headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/70.0.3538.110 Safari/537.36'} timeout = 5 gzh_name = 'pytest文檔' My_GZH = WeixinSpider(gzh_name,5,'pytest') My_GZH.get_page_url() # print(My_GZH.page_url) My_GZH.get_article_url() # print(My_GZH.article_list) for article in My_GZH.article_list: for (key,value) in article.items(): url=value html_response = requests.get(url,headers=headers,timeout=timeout) myHTMLParser = MyHTMLParser(key) myHTMLParser.feed(html_response.text) myHTMLParser.doc.save(myHTMLParser.docfile)
HTML2doc.py
from html.parser import HTMLParser import requests from docx import Document import re from docx.shared import RGBColor import docx class MyHTMLParser(HTMLParser): def __init__(self,docname): HTMLParser.__init__(self) self.docname=docname self.docfile = r"D:\pytest\%s.doc"%self.docname self.doc=Document() self.title = False self.code = False self.text='' self.processing =None self.codeprocessing =None self.picindex = 1 self.headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/70.0.3538.110 Safari/537.36'} self.timeout = 5 def handle_startendtag(self, tag, attrs): # 圖片的處理比較復(fù)雜,首先需要找到對(duì)應(yīng)的圖片的url,然后下載并寫入doc中 if tag == "img": if len(attrs) == 0: pass else: for (variable, value) in attrs: if variable == "data-type": picname = r"D:\pytest\%s%s.%s" % (self.docname, self.picindex, value) # print(picname) if variable == "data-src": picdata = requests.get(value, headers=self.headers, timeout=self.timeout) # print(value) self.picindex = self.picindex + 1 # print(self.picindex) with open(picname, "wb") as pic: pic.write(picdata.content) try: self.doc.add_picture(picname) except docx.image.exceptions.UnexpectedEndOfFileError as e: print(e) def handle_starttag(self, tag, attrs): if re.match(r"h(\d)", tag): self.title = True if tag =="p": self.processing = tag if tag == "code": self.code = True self.codeprocessing = tag def handle_data(self, data): if self.title == True: self.doc.add_heading(data, level=2) # if self.in_div == True and self.tag == "p": if self.processing: self.text = self.text + data if self.code == True: p =self.doc.add_paragraph() run=p.add_run(data) run.font.color.rgb = RGBColor(111,111,111) def handle_endtag(self, tag): self.title = False # self.code = False if tag == self.processing: self.doc.add_paragraph(self.text) self.processing = None self.text='' if tag == self.codeprocessing: self.code =False
運(yùn)行結(jié)果:
缺少部分文檔,如pytest文檔4,是因?yàn)樗压肺⑿盼恼滤阉鹘Y(jié)果中就沒有
以上就是本文的全部?jī)?nèi)容,希望對(duì)大家的學(xué)習(xí)有所幫助,也希望大家多多支持腳本之家。
相關(guān)文章
詳解Python odoo中嵌入html簡(jiǎn)單的分頁(yè)功能
在odoo中,通過iframe嵌入 html,頁(yè)面數(shù)據(jù)則通過controllers獲取,使用jinja2模板傳值渲染。這篇文章主要介紹了Python odoo中嵌入html簡(jiǎn)單的分頁(yè)功能 ,需要的朋友可以參考下2019-05-05TensorFlow 2.0之后動(dòng)態(tài)分配顯存方式
這篇文章主要介紹了TensorFlow 2.0之后動(dòng)態(tài)分配顯存方式,具有很好的參考價(jià)值,希望對(duì)大家有所幫助。如有錯(cuò)誤或未考慮完全的地方,望不吝賜教2022-12-12Django REST framework內(nèi)置路由用法
這篇文章主要介紹了Django REST framework內(nèi)置路由用法,文中通過示例代碼介紹的非常詳細(xì),對(duì)大家的學(xué)習(xí)或者工作具有一定的參考學(xué)習(xí)價(jià)值,需要的朋友可以參考下2019-07-07淺談selenium如何應(yīng)對(duì)網(wǎng)頁(yè)內(nèi)容需要鼠標(biāo)滾動(dòng)加載的問題
這篇文章主要介紹了淺談selenium如何應(yīng)對(duì)網(wǎng)頁(yè)內(nèi)容需要鼠標(biāo)滾動(dòng)加載的問題,具有很好的參考價(jià)值,希望對(duì)大家有所幫助。一起跟隨小編過來(lái)看看吧2020-03-03python測(cè)試攻略pytest.main()隱藏利器實(shí)例探究
在Pytest測(cè)試框架中,pytest.main()是一個(gè)重要的功能,用于啟動(dòng)測(cè)試執(zhí)行,它允許以不同方式運(yùn)行測(cè)試,傳遞參數(shù)和配置選項(xiàng),本文將深入探討pytest.main()的核心功能,提供豐富的示例代碼和更全面的內(nèi)容,2024-01-01python 簡(jiǎn)易計(jì)算器程序,代碼就幾行
運(yùn)行環(huán)境:python 3.1,代碼比較短,大家可以參考下。2009-08-08