Dev (Sourcery refactored) - #18

Open
sourcery-ai[bot] wants to merge 1 commit into
devfrom
sourcery/dev
Open

Dev (Sourcery refactored)#18
sourcery-ai[bot] wants to merge 1 commit into
devfrom
sourcery/dev

Conversation

@sourcery-ai

Copy link
Copy Markdown

Pull Request #17 refactored by Sourcery.

If you're happy with these changes, merge this Pull Request using the Squash and merge strategy.

NOTE: As code is pushed to the original Pull Request, Sourcery will
re-run and update (force-push) this Pull Request with new refactorings as
necessary. If Sourcery finds no refactorings at any point, this Pull Request
will be closed automatically.

See our documentation here.

Run Sourcery locally

Reduce the feedback loop during development by using the Sourcery editor plugin:

Review changes via command line

To manually merge these changes, make sure you're on the dev branch, then run:

git fetch origin sourcery/dev
git merge --ff-only FETCH_HEAD
git reset HEAD^

Help us improve this pull request!

@sourcery-aisourcery-aiBot left a comment

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Due to GitHub API limits, only the first 60 comments can be shown.

Comment threadSchedule/demo4.py

def greet(name):
print('Hello {}'.format(name))
print(f'Hello {name}')

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function greet refactored with the following changes:

Comment threadSchedule/demo8.py
"""
def job1():
print("I'm running on threads %s" % threading.current_thread())
print(f"I'm running on threads {threading.current_thread()}")

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function job1 refactored with the following changes:

Comment threadSchedule/demo8.py
print(f"I'm running on threads {threading.current_thread()}")
def job2():
print("I'm running on threads %s" % threading.current_thread())
print(f"I'm running on threads {threading.current_thread()}")

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function job2 refactored with the following changes:

Comment threadSchedule/demo8.py
print(f"I'm running on threads {threading.current_thread()}")
def job3():
print("I'm running on threads %s" % threading.current_thread())
print(f"I'm running on threads {threading.current_thread()}")

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function job3 refactored with the following changes:

Comment on lines -13 to +15
CHROME_DRIVER_MAPPING_FILE = r"{}\mapping.json".format(CHROME_DRIVER_FOLDER)
CHROME_DRIVER_EXE = r"{}\chromedriver.exe".format(CHROME_DRIVER_FOLDER)
CHROME_DRIVER_ZIP = r"{}\chromedriver_win32.zip".format(CHROME_DRIVER_FOLDER)
CHROME_DRIVER_MAPPING_FILE = f"{CHROME_DRIVER_FOLDER}\mapping.json"
CHROME_DRIVER_EXE = f"{CHROME_DRIVER_FOLDER}\chromedriver.exe"
CHROME_DRIVER_ZIP = f"{CHROME_DRIVER_FOLDER}\chromedriver_win32.zip"

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Lines 13-15 refactored with the following changes:

Comment on lines -386 to +371
ret = []
for i in range(1, page):
ret.append('http://www.samair.ru/proxy/proxy-%(num)02d.htm' % {'num': i})
return ret
return [
'http://www.samair.ru/proxy/proxy-%(num)02d.htm' % {'num': i}
for i in range(1, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_3 refactored with the following changes:

Comment on lines -406 to +388
ip = match[0] + "." + match[1] + match[2]
ip = f"{match[0]}.{match[1]}{match[2]}"

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function parse_page_3 refactored with the following changes:

Comment on lines -432 to +417
ret = []
for i in range(1, page):
ret.append('http://www.pass-e.com/proxy/index.php?page=%(n)01d' % {'n': i})
return ret
return [
'http://www.pass-e.com/proxy/index.php?page=%(n)01d' % {'n': i}
for i in range(1, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_4 refactored with the following changes:

Comment on lines -476 to +461
ret = []
for i in range(1, page):
ret.append('http://www.ipfree.cn/index2.asp?page=%(num)01d' % {'num': i})
return ret
return [
'http://www.ipfree.cn/index2.asp?page=%(num)01d' % {'num': i}
for i in range(1, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_5 refactored with the following changes:

Comment on lines -508 to +493
ret = []
for i in range(1, page):
ret.append('http://www.cnproxy.com/proxy%(num)01d.html' % {'num': i})
return ret
return [
'http://www.cnproxy.com/proxy%(num)01d.html' % {'num': i}
for i in range(1, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_6 refactored with the following changes:

Comment on lines +507 to -528
type = -1 # 该网站未提供代理服务器类型
for match in matches:
ip = match[0]
port = match[1]
type = -1 # 该网站未提供代理服务器类型

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function parse_page_6 refactored with the following changes:

Comment on lines -595 to +580
ret = []
for i in range(0, page):
ret.append('http://proxylist.sakura.ne.jp/index.htm?pages=%(n)01d' % {'n': i})
return ret
return [
'http://proxylist.sakura.ne.jp/index.htm?pages=%(n)01d' % {'n': i}
for i in range(0, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_9 refactored with the following changes:

Comment on lines -616 to +598
if (type == 'Anonymous'):
type = 1
else:
type = -1
type = 1 if (type == 'Anonymous') else -1

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function parse_page_9 refactored with the following changes:

Comment on lines -634 to +616
ret = []
for i in range(1, page):
ret.append('http://www.publicproxyservers.com/page%(n)01d.html' % {'n': i})
return ret
return [
'http://www.publicproxyservers.com/page%(n)01d.html' % {'n': i}
for i in range(1, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_10 refactored with the following changes:

Comment on lines -678 to +667
ret = []
for i in range(1, page):
ret.append('http://www.my-proxy.com/list/proxy.php?list=%(n)01d' % {'n': i})

ret.append('http://www.my-proxy.com/list/proxy.php?list=s1')
ret.append('http://www.my-proxy.com/list/proxy.php?list=s2')
ret.append('http://www.my-proxy.com/list/proxy.php?list=s3')
ret = [
'http://www.my-proxy.com/list/proxy.php?list=%(n)01d' % {'n': i}
for i in range(1, page)
]
ret.extend(
(
'http://www.my-proxy.com/list/proxy.php?list=s1',
'http://www.my-proxy.com/list/proxy.php?list=s2',
'http://www.my-proxy.com/list/proxy.php?list=s3',
)
)

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_11 refactored with the following changes:

urls = []
for page in range(beg, end):
urls.append('http://www.baidu.com?&page=%d' % page)
urls = ['http://www.baidu.com?&page=%d' % page for page in range(beg, end)]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function run_spider refactored with the following changes:

Comment on lines -92 to +90
res = []
for i in range(len(s) - 1):
res.append((s[i], s[i + 1]))
res = [(s[i], s[i + 1]) for i in range(len(s) - 1)]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function main refactored with the following changes:

Comment on lines -64 to -68
proxies = {
'http': '{proxy_type}://{ip}:{port}'.format(proxy_type=proxy_type, ip=ip, port=port),
'https': '{proxy_type}://{ip}:{port}'.format(proxy_type=proxy_type, ip=ip, port=port),
return {
'http': '{proxy_type}://{ip}:{port}'.format(
proxy_type=proxy_type, ip=ip, port=port
),
'https': '{proxy_type}://{ip}:{port}'.format(
proxy_type=proxy_type, ip=ip, port=port
),
}
return proxies

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function get_proxy_dict refactored with the following changes:

urls = []
for page in range(1, 1000):
urls.append('http://www.jb51.net/article/%s.htm' % page)
urls = [f'http://www.jb51.net/article/{page}.htm' for page in range(1, 1000)]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function main refactored with the following changes:

Comment threadcrawler/src/test.py
class TestSyncSpider(SyncSpider):
def handle_html(self, url, html):
print(html)
pass

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function TestSyncSpider.handle_html refactored with the following changes:

Comment threadcrawler/src/test.py
class TestAsyncSpider(AsyncSpider):
def handle_html(self, url, html):
print(html)
pass

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function TestAsyncSpider.handle_html refactored with the following changes:

Comment threadcrawler/src/test.py
Comment on lines -24 to +25
urls = []
for page in range(1, 1000):
#urls.append('http://www.jb51.net/article/%s.htm' % page)
urls.append('http://www.imooc.com/data/check_f.php?page=%d'%page)
urls = [
'http://www.imooc.com/data/check_f.php?page=%d' % page
for page in range(1, 1000)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Lines 24-27 refactored with the following changes:

This removes the following comments ( why? ):

#urls.append('http://www.jb51.net/article/%s.htm' % page)

Comment on lines -64 to +69
func_name = 'get_' + field
func_name = f'get_{field}'
xpath_str = self.xpath_dict.get(field)
if hasattr(self, func_name):
return getattr(self, func_name)(xpath_str)
else:
self.logger.debug(field, self.url)
return self.parser.xpath(xpath_str)[0].strip() if xpath_str else ''
self.logger.debug(field, self.url)
return self.parser.xpath(xpath_str)[0].strip() if xpath_str else ''

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function XpathCrawler.get_field refactored with the following changes:

Comment on lines -73 to +75
xpath_result = {}
for field, xpath_string in self.xpath_dict.items():
xpath_result[field] = makes(self.get_field(field)) # to utf8
return xpath_result
return {
field: makes(self.get_field(field))
for field, xpath_string in self.xpath_dict.items()
}

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function XpathCrawler.get_result refactored with the following changes:

This removes the following comments ( why? ):

# to utf8

Comment on lines -93 to +92
category_urls = []
for href in category_hrefs:
category_urls.append(urljoin(self.domain, href))
category_urls = [urljoin(self.domain, href) for href in category_hrefs]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function FantasyhairbuySite.generate_category_urls refactored with the following changes:

Comment on lines -11 to +14
backup_name = os.path.basename(file_path) + '_' + datetime.datetime.now().strftime('%Y%m%d%H%M%S')
backup_name = (
f'{os.path.basename(file_path)}_'
+ datetime.datetime.now().strftime('%Y%m%d%H%M%S')
)

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function backup_file refactored with the following changes:

Comment on lines -44 to +47
backup_list = glob.glob(os.path.join(path, pattern + '_*'))
backup_list = glob.glob(os.path.join(path, f'{pattern}_*'))

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function main refactored with the following changes:


# if it's anything else, return it in its original form
return data
return data.encode('utf-8') if str(type(data)) == "<type 'unicode'>" else data

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function _byteify refactored with the following changes:

This removes the following comments ( why? ):

# if it's anything else, return it in its original form

try:
res = query.find()
return res
return query.find()

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function LeanCloudApi.get_skip_obj_list refactored with the following changes:

Comment on lines -49 to +48
img_info_url = img_url + '?imageInfo'
img_info_url = f'{img_url}?imageInfo'

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function LeanCloudApi.add_img_info refactored with the following changes:

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Dev (Sourcery refactored) - #18

Open
sourcery-ai[bot] wants to merge 1 commit into
devfrom
sourcery/dev
Open

Dev (Sourcery refactored)#18
sourcery-ai[bot] wants to merge 1 commit into
devfrom
sourcery/dev

Conversation

@sourcery-ai

Copy link
Copy Markdown

Pull Request #17 refactored by Sourcery.

If you're happy with these changes, merge this Pull Request using the Squash and merge strategy.

NOTE: As code is pushed to the original Pull Request, Sourcery will
re-run and update (force-push) this Pull Request with new refactorings as
necessary. If Sourcery finds no refactorings at any point, this Pull Request
will be closed automatically.

See our documentation here.

Run Sourcery locally

Reduce the feedback loop during development by using the Sourcery editor plugin:

Review changes via command line

To manually merge these changes, make sure you're on the dev branch, then run:

git fetch origin sourcery/dev
git merge --ff-only FETCH_HEAD
git reset HEAD^

Help us improve this pull request!

@sourcery-aisourcery-aiBot left a comment

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Due to GitHub API limits, only the first 60 comments can be shown.

Comment threadSchedule/demo4.py

def greet(name):
print('Hello {}'.format(name))
print(f'Hello {name}')

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function greet refactored with the following changes:

Comment threadSchedule/demo8.py
"""
def job1():
print("I'm running on threads %s" % threading.current_thread())
print(f"I'm running on threads {threading.current_thread()}")

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function job1 refactored with the following changes:

Comment threadSchedule/demo8.py
print(f"I'm running on threads {threading.current_thread()}")
def job2():
print("I'm running on threads %s" % threading.current_thread())
print(f"I'm running on threads {threading.current_thread()}")

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function job2 refactored with the following changes:

Comment threadSchedule/demo8.py
print(f"I'm running on threads {threading.current_thread()}")
def job3():
print("I'm running on threads %s" % threading.current_thread())
print(f"I'm running on threads {threading.current_thread()}")

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function job3 refactored with the following changes:

Comment on lines -13 to +15
CHROME_DRIVER_MAPPING_FILE = r"{}\mapping.json".format(CHROME_DRIVER_FOLDER)
CHROME_DRIVER_EXE = r"{}\chromedriver.exe".format(CHROME_DRIVER_FOLDER)
CHROME_DRIVER_ZIP = r"{}\chromedriver_win32.zip".format(CHROME_DRIVER_FOLDER)
CHROME_DRIVER_MAPPING_FILE = f"{CHROME_DRIVER_FOLDER}\mapping.json"
CHROME_DRIVER_EXE = f"{CHROME_DRIVER_FOLDER}\chromedriver.exe"
CHROME_DRIVER_ZIP = f"{CHROME_DRIVER_FOLDER}\chromedriver_win32.zip"

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Lines 13-15 refactored with the following changes:

Comment on lines -386 to +371
ret = []
for i in range(1, page):
ret.append('http://www.samair.ru/proxy/proxy-%(num)02d.htm' % {'num': i})
return ret
return [
'http://www.samair.ru/proxy/proxy-%(num)02d.htm' % {'num': i}
for i in range(1, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_3 refactored with the following changes:

Comment on lines -406 to +388
ip = match[0] + "." + match[1] + match[2]
ip = f"{match[0]}.{match[1]}{match[2]}"

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function parse_page_3 refactored with the following changes:

Comment on lines -432 to +417
ret = []
for i in range(1, page):
ret.append('http://www.pass-e.com/proxy/index.php?page=%(n)01d' % {'n': i})
return ret
return [
'http://www.pass-e.com/proxy/index.php?page=%(n)01d' % {'n': i}
for i in range(1, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_4 refactored with the following changes:

Comment on lines -476 to +461
ret = []
for i in range(1, page):
ret.append('http://www.ipfree.cn/index2.asp?page=%(num)01d' % {'num': i})
return ret
return [
'http://www.ipfree.cn/index2.asp?page=%(num)01d' % {'num': i}
for i in range(1, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_5 refactored with the following changes:

Comment on lines -508 to +493
ret = []
for i in range(1, page):
ret.append('http://www.cnproxy.com/proxy%(num)01d.html' % {'num': i})
return ret
return [
'http://www.cnproxy.com/proxy%(num)01d.html' % {'num': i}
for i in range(1, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_6 refactored with the following changes:

Comment on lines +507 to -528
type = -1 # 该网站未提供代理服务器类型
for match in matches:
ip = match[0]
port = match[1]
type = -1 # 该网站未提供代理服务器类型

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function parse_page_6 refactored with the following changes:

Comment on lines -595 to +580
ret = []
for i in range(0, page):
ret.append('http://proxylist.sakura.ne.jp/index.htm?pages=%(n)01d' % {'n': i})
return ret
return [
'http://proxylist.sakura.ne.jp/index.htm?pages=%(n)01d' % {'n': i}
for i in range(0, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_9 refactored with the following changes:

Comment on lines -616 to +598
if (type == 'Anonymous'):
type = 1
else:
type = -1
type = 1 if (type == 'Anonymous') else -1

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function parse_page_9 refactored with the following changes:

Comment on lines -634 to +616
ret = []
for i in range(1, page):
ret.append('http://www.publicproxyservers.com/page%(n)01d.html' % {'n': i})
return ret
return [
'http://www.publicproxyservers.com/page%(n)01d.html' % {'n': i}
for i in range(1, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_10 refactored with the following changes:

Comment on lines -678 to +667
ret = []
for i in range(1, page):
ret.append('http://www.my-proxy.com/list/proxy.php?list=%(n)01d' % {'n': i})

ret.append('http://www.my-proxy.com/list/proxy.php?list=s1')
ret.append('http://www.my-proxy.com/list/proxy.php?list=s2')
ret.append('http://www.my-proxy.com/list/proxy.php?list=s3')
ret = [
'http://www.my-proxy.com/list/proxy.php?list=%(n)01d' % {'n': i}
for i in range(1, page)
]
ret.extend(
(
'http://www.my-proxy.com/list/proxy.php?list=s1',
'http://www.my-proxy.com/list/proxy.php?list=s2',
'http://www.my-proxy.com/list/proxy.php?list=s3',
)
)

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_11 refactored with the following changes:

urls = []
for page in range(beg, end):
urls.append('http://www.baidu.com?&page=%d' % page)
urls = ['http://www.baidu.com?&page=%d' % page for page in range(beg, end)]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function run_spider refactored with the following changes:

Comment on lines -92 to +90
res = []
for i in range(len(s) - 1):
res.append((s[i], s[i + 1]))
res = [(s[i], s[i + 1]) for i in range(len(s) - 1)]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function main refactored with the following changes:

Comment on lines -64 to -68
proxies = {
'http': '{proxy_type}://{ip}:{port}'.format(proxy_type=proxy_type, ip=ip, port=port),
'https': '{proxy_type}://{ip}:{port}'.format(proxy_type=proxy_type, ip=ip, port=port),
return {
'http': '{proxy_type}://{ip}:{port}'.format(
proxy_type=proxy_type, ip=ip, port=port
),
'https': '{proxy_type}://{ip}:{port}'.format(
proxy_type=proxy_type, ip=ip, port=port
),
}
return proxies

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function get_proxy_dict refactored with the following changes:

urls = []
for page in range(1, 1000):
urls.append('http://www.jb51.net/article/%s.htm' % page)
urls = [f'http://www.jb51.net/article/{page}.htm' for page in range(1, 1000)]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function main refactored with the following changes:

Comment threadcrawler/src/test.py
class TestSyncSpider(SyncSpider):
def handle_html(self, url, html):
print(html)
pass

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function TestSyncSpider.handle_html refactored with the following changes:

Comment threadcrawler/src/test.py
class TestAsyncSpider(AsyncSpider):
def handle_html(self, url, html):
print(html)
pass

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function TestAsyncSpider.handle_html refactored with the following changes:

Comment threadcrawler/src/test.py
Comment on lines -24 to +25
urls = []
for page in range(1, 1000):
#urls.append('http://www.jb51.net/article/%s.htm' % page)
urls.append('http://www.imooc.com/data/check_f.php?page=%d'%page)
urls = [
'http://www.imooc.com/data/check_f.php?page=%d' % page
for page in range(1, 1000)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Lines 24-27 refactored with the following changes:

This removes the following comments ( why? ):

#urls.append('http://www.jb51.net/article/%s.htm' % page)

Comment on lines -64 to +69
func_name = 'get_' + field
func_name = f'get_{field}'
xpath_str = self.xpath_dict.get(field)
if hasattr(self, func_name):
return getattr(self, func_name)(xpath_str)
else:
self.logger.debug(field, self.url)
return self.parser.xpath(xpath_str)[0].strip() if xpath_str else ''
self.logger.debug(field, self.url)
return self.parser.xpath(xpath_str)[0].strip() if xpath_str else ''

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function XpathCrawler.get_field refactored with the following changes:

Comment on lines -73 to +75
xpath_result = {}
for field, xpath_string in self.xpath_dict.items():
xpath_result[field] = makes(self.get_field(field)) # to utf8
return xpath_result
return {
field: makes(self.get_field(field))
for field, xpath_string in self.xpath_dict.items()
}

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function XpathCrawler.get_result refactored with the following changes:

This removes the following comments ( why? ):

# to utf8

Comment on lines -93 to +92
category_urls = []
for href in category_hrefs:
category_urls.append(urljoin(self.domain, href))
category_urls = [urljoin(self.domain, href) for href in category_hrefs]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function FantasyhairbuySite.generate_category_urls refactored with the following changes:

Comment on lines -11 to +14
backup_name = os.path.basename(file_path) + '_' + datetime.datetime.now().strftime('%Y%m%d%H%M%S')
backup_name = (
f'{os.path.basename(file_path)}_'
+ datetime.datetime.now().strftime('%Y%m%d%H%M%S')
)

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function backup_file refactored with the following changes:

Comment on lines -44 to +47
backup_list = glob.glob(os.path.join(path, pattern + '_*'))
backup_list = glob.glob(os.path.join(path, f'{pattern}_*'))

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function main refactored with the following changes:


# if it's anything else, return it in its original form
return data
return data.encode('utf-8') if str(type(data)) == "<type 'unicode'>" else data

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function _byteify refactored with the following changes:

This removes the following comments ( why? ):

# if it's anything else, return it in its original form

try:
res = query.find()
return res
return query.find()

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function LeanCloudApi.get_skip_obj_list refactored with the following changes:

Comment on lines -49 to +48
img_info_url = img_url + '?imageInfo'
img_info_url = f'{img_url}?imageInfo'

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function LeanCloudApi.add_img_info refactored with the following changes:

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Dev (Sourcery refactored) - #18

Open
sourcery-ai[bot] wants to merge 1 commit into
devfrom
sourcery/dev
Open

Dev (Sourcery refactored)#18
sourcery-ai[bot] wants to merge 1 commit into
devfrom
sourcery/dev

Conversation

@sourcery-ai

Copy link
Copy Markdown

Pull Request #17 refactored by Sourcery.

If you're happy with these changes, merge this Pull Request using the Squash and merge strategy.

NOTE: As code is pushed to the original Pull Request, Sourcery will
re-run and update (force-push) this Pull Request with new refactorings as
necessary. If Sourcery finds no refactorings at any point, this Pull Request
will be closed automatically.

See our documentation here.

Run Sourcery locally

Reduce the feedback loop during development by using the Sourcery editor plugin:

Review changes via command line

To manually merge these changes, make sure you're on the dev branch, then run:

git fetch origin sourcery/dev
git merge --ff-only FETCH_HEAD
git reset HEAD^

Help us improve this pull request!

@sourcery-aisourcery-aiBot left a comment

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Due to GitHub API limits, only the first 60 comments can be shown.

Comment threadSchedule/demo4.py

def greet(name):
print('Hello {}'.format(name))
print(f'Hello {name}')

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function greet refactored with the following changes:

Comment threadSchedule/demo8.py
"""
def job1():
print("I'm running on threads %s" % threading.current_thread())
print(f"I'm running on threads {threading.current_thread()}")

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function job1 refactored with the following changes:

Comment threadSchedule/demo8.py
print(f"I'm running on threads {threading.current_thread()}")
def job2():
print("I'm running on threads %s" % threading.current_thread())
print(f"I'm running on threads {threading.current_thread()}")

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function job2 refactored with the following changes:

Comment threadSchedule/demo8.py
print(f"I'm running on threads {threading.current_thread()}")
def job3():
print("I'm running on threads %s" % threading.current_thread())
print(f"I'm running on threads {threading.current_thread()}")

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function job3 refactored with the following changes:

Comment on lines -13 to +15
CHROME_DRIVER_MAPPING_FILE = r"{}\mapping.json".format(CHROME_DRIVER_FOLDER)
CHROME_DRIVER_EXE = r"{}\chromedriver.exe".format(CHROME_DRIVER_FOLDER)
CHROME_DRIVER_ZIP = r"{}\chromedriver_win32.zip".format(CHROME_DRIVER_FOLDER)
CHROME_DRIVER_MAPPING_FILE = f"{CHROME_DRIVER_FOLDER}\mapping.json"
CHROME_DRIVER_EXE = f"{CHROME_DRIVER_FOLDER}\chromedriver.exe"
CHROME_DRIVER_ZIP = f"{CHROME_DRIVER_FOLDER}\chromedriver_win32.zip"

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Lines 13-15 refactored with the following changes:

Comment on lines -386 to +371
ret = []
for i in range(1, page):
ret.append('http://www.samair.ru/proxy/proxy-%(num)02d.htm' % {'num': i})
return ret
return [
'http://www.samair.ru/proxy/proxy-%(num)02d.htm' % {'num': i}
for i in range(1, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_3 refactored with the following changes:

Comment on lines -406 to +388
ip = match[0] + "." + match[1] + match[2]
ip = f"{match[0]}.{match[1]}{match[2]}"

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function parse_page_3 refactored with the following changes:

Comment on lines -432 to +417
ret = []
for i in range(1, page):
ret.append('http://www.pass-e.com/proxy/index.php?page=%(n)01d' % {'n': i})
return ret
return [
'http://www.pass-e.com/proxy/index.php?page=%(n)01d' % {'n': i}
for i in range(1, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_4 refactored with the following changes:

Comment on lines -476 to +461
ret = []
for i in range(1, page):
ret.append('http://www.ipfree.cn/index2.asp?page=%(num)01d' % {'num': i})
return ret
return [
'http://www.ipfree.cn/index2.asp?page=%(num)01d' % {'num': i}
for i in range(1, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_5 refactored with the following changes:

Comment on lines -508 to +493
ret = []
for i in range(1, page):
ret.append('http://www.cnproxy.com/proxy%(num)01d.html' % {'num': i})
return ret
return [
'http://www.cnproxy.com/proxy%(num)01d.html' % {'num': i}
for i in range(1, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_6 refactored with the following changes:

Comment on lines +507 to -528
type = -1 # 该网站未提供代理服务器类型
for match in matches:
ip = match[0]
port = match[1]
type = -1 # 该网站未提供代理服务器类型

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function parse_page_6 refactored with the following changes:

Comment on lines -595 to +580
ret = []
for i in range(0, page):
ret.append('http://proxylist.sakura.ne.jp/index.htm?pages=%(n)01d' % {'n': i})
return ret
return [
'http://proxylist.sakura.ne.jp/index.htm?pages=%(n)01d' % {'n': i}
for i in range(0, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_9 refactored with the following changes:

Comment on lines -616 to +598
if (type == 'Anonymous'):
type = 1
else:
type = -1
type = 1 if (type == 'Anonymous') else -1

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function parse_page_9 refactored with the following changes:

Comment on lines -634 to +616
ret = []
for i in range(1, page):
ret.append('http://www.publicproxyservers.com/page%(n)01d.html' % {'n': i})
return ret
return [
'http://www.publicproxyservers.com/page%(n)01d.html' % {'n': i}
for i in range(1, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_10 refactored with the following changes:

Comment on lines -678 to +667
ret = []
for i in range(1, page):
ret.append('http://www.my-proxy.com/list/proxy.php?list=%(n)01d' % {'n': i})

ret.append('http://www.my-proxy.com/list/proxy.php?list=s1')
ret.append('http://www.my-proxy.com/list/proxy.php?list=s2')
ret.append('http://www.my-proxy.com/list/proxy.php?list=s3')
ret = [
'http://www.my-proxy.com/list/proxy.php?list=%(n)01d' % {'n': i}
for i in range(1, page)
]
ret.extend(
(
'http://www.my-proxy.com/list/proxy.php?list=s1',
'http://www.my-proxy.com/list/proxy.php?list=s2',
'http://www.my-proxy.com/list/proxy.php?list=s3',
)
)

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_11 refactored with the following changes:

urls = []
for page in range(beg, end):
urls.append('http://www.baidu.com?&page=%d' % page)
urls = ['http://www.baidu.com?&page=%d' % page for page in range(beg, end)]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function run_spider refactored with the following changes:

Comment on lines -92 to +90
res = []
for i in range(len(s) - 1):
res.append((s[i], s[i + 1]))
res = [(s[i], s[i + 1]) for i in range(len(s) - 1)]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function main refactored with the following changes:

Comment on lines -64 to -68
proxies = {
'http': '{proxy_type}://{ip}:{port}'.format(proxy_type=proxy_type, ip=ip, port=port),
'https': '{proxy_type}://{ip}:{port}'.format(proxy_type=proxy_type, ip=ip, port=port),
return {
'http': '{proxy_type}://{ip}:{port}'.format(
proxy_type=proxy_type, ip=ip, port=port
),
'https': '{proxy_type}://{ip}:{port}'.format(
proxy_type=proxy_type, ip=ip, port=port
),
}
return proxies

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function get_proxy_dict refactored with the following changes:

urls = []
for page in range(1, 1000):
urls.append('http://www.jb51.net/article/%s.htm' % page)
urls = [f'http://www.jb51.net/article/{page}.htm' for page in range(1, 1000)]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function main refactored with the following changes:

Comment threadcrawler/src/test.py
class TestSyncSpider(SyncSpider):
def handle_html(self, url, html):
print(html)
pass

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function TestSyncSpider.handle_html refactored with the following changes:

Comment threadcrawler/src/test.py
class TestAsyncSpider(AsyncSpider):
def handle_html(self, url, html):
print(html)
pass

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function TestAsyncSpider.handle_html refactored with the following changes:

Comment threadcrawler/src/test.py
Comment on lines -24 to +25
urls = []
for page in range(1, 1000):
#urls.append('http://www.jb51.net/article/%s.htm' % page)
urls.append('http://www.imooc.com/data/check_f.php?page=%d'%page)
urls = [
'http://www.imooc.com/data/check_f.php?page=%d' % page
for page in range(1, 1000)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Lines 24-27 refactored with the following changes:

This removes the following comments ( why? ):

#urls.append('http://www.jb51.net/article/%s.htm' % page)

Comment on lines -64 to +69
func_name = 'get_' + field
func_name = f'get_{field}'
xpath_str = self.xpath_dict.get(field)
if hasattr(self, func_name):
return getattr(self, func_name)(xpath_str)
else:
self.logger.debug(field, self.url)
return self.parser.xpath(xpath_str)[0].strip() if xpath_str else ''
self.logger.debug(field, self.url)
return self.parser.xpath(xpath_str)[0].strip() if xpath_str else ''

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function XpathCrawler.get_field refactored with the following changes:

Comment on lines -73 to +75
xpath_result = {}
for field, xpath_string in self.xpath_dict.items():
xpath_result[field] = makes(self.get_field(field)) # to utf8
return xpath_result
return {
field: makes(self.get_field(field))
for field, xpath_string in self.xpath_dict.items()
}

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function XpathCrawler.get_result refactored with the following changes:

This removes the following comments ( why? ):

# to utf8

Comment on lines -93 to +92
category_urls = []
for href in category_hrefs:
category_urls.append(urljoin(self.domain, href))
category_urls = [urljoin(self.domain, href) for href in category_hrefs]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function FantasyhairbuySite.generate_category_urls refactored with the following changes:

Comment on lines -11 to +14
backup_name = os.path.basename(file_path) + '_' + datetime.datetime.now().strftime('%Y%m%d%H%M%S')
backup_name = (
f'{os.path.basename(file_path)}_'
+ datetime.datetime.now().strftime('%Y%m%d%H%M%S')
)

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function backup_file refactored with the following changes:

Comment on lines -44 to +47
backup_list = glob.glob(os.path.join(path, pattern + '_*'))
backup_list = glob.glob(os.path.join(path, f'{pattern}_*'))

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function main refactored with the following changes:


# if it's anything else, return it in its original form
return data
return data.encode('utf-8') if str(type(data)) == "<type 'unicode'>" else data

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function _byteify refactored with the following changes:

This removes the following comments ( why? ):

# if it's anything else, return it in its original form

try:
res = query.find()
return res
return query.find()

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function LeanCloudApi.get_skip_obj_list refactored with the following changes:

Comment on lines -49 to +48
img_info_url = img_url + '?imageInfo'
img_info_url = f'{img_url}?imageInfo'

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function LeanCloudApi.add_img_info refactored with the following changes:

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Dev (Sourcery refactored) - #18

Open
sourcery-ai[bot] wants to merge 1 commit into
devfrom
sourcery/dev
Open

Dev (Sourcery refactored)#18
sourcery-ai[bot] wants to merge 1 commit into
devfrom
sourcery/dev

Conversation

@sourcery-ai

Copy link
Copy Markdown

Pull Request #17 refactored by Sourcery.

If you're happy with these changes, merge this Pull Request using the Squash and merge strategy.

NOTE: As code is pushed to the original Pull Request, Sourcery will
re-run and update (force-push) this Pull Request with new refactorings as
necessary. If Sourcery finds no refactorings at any point, this Pull Request
will be closed automatically.

See our documentation here.

Run Sourcery locally

Reduce the feedback loop during development by using the Sourcery editor plugin:

Review changes via command line

To manually merge these changes, make sure you're on the dev branch, then run:

git fetch origin sourcery/dev
git merge --ff-only FETCH_HEAD
git reset HEAD^

Help us improve this pull request!

@sourcery-aisourcery-aiBot left a comment

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Due to GitHub API limits, only the first 60 comments can be shown.

Comment threadSchedule/demo4.py

def greet(name):
print('Hello {}'.format(name))
print(f'Hello {name}')

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function greet refactored with the following changes:

Comment threadSchedule/demo8.py
"""
def job1():
print("I'm running on threads %s" % threading.current_thread())
print(f"I'm running on threads {threading.current_thread()}")

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function job1 refactored with the following changes:

Comment threadSchedule/demo8.py
print(f"I'm running on threads {threading.current_thread()}")
def job2():
print("I'm running on threads %s" % threading.current_thread())
print(f"I'm running on threads {threading.current_thread()}")

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function job2 refactored with the following changes:

Comment threadSchedule/demo8.py
print(f"I'm running on threads {threading.current_thread()}")
def job3():
print("I'm running on threads %s" % threading.current_thread())
print(f"I'm running on threads {threading.current_thread()}")

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function job3 refactored with the following changes:

Comment on lines -13 to +15
CHROME_DRIVER_MAPPING_FILE = r"{}\mapping.json".format(CHROME_DRIVER_FOLDER)
CHROME_DRIVER_EXE = r"{}\chromedriver.exe".format(CHROME_DRIVER_FOLDER)
CHROME_DRIVER_ZIP = r"{}\chromedriver_win32.zip".format(CHROME_DRIVER_FOLDER)
CHROME_DRIVER_MAPPING_FILE = f"{CHROME_DRIVER_FOLDER}\mapping.json"
CHROME_DRIVER_EXE = f"{CHROME_DRIVER_FOLDER}\chromedriver.exe"
CHROME_DRIVER_ZIP = f"{CHROME_DRIVER_FOLDER}\chromedriver_win32.zip"

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Lines 13-15 refactored with the following changes:

Comment on lines -386 to +371
ret = []
for i in range(1, page):
ret.append('http://www.samair.ru/proxy/proxy-%(num)02d.htm' % {'num': i})
return ret
return [
'http://www.samair.ru/proxy/proxy-%(num)02d.htm' % {'num': i}
for i in range(1, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_3 refactored with the following changes:

Comment on lines -406 to +388
ip = match[0] + "." + match[1] + match[2]
ip = f"{match[0]}.{match[1]}{match[2]}"

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function parse_page_3 refactored with the following changes:

Comment on lines -432 to +417
ret = []
for i in range(1, page):
ret.append('http://www.pass-e.com/proxy/index.php?page=%(n)01d' % {'n': i})
return ret
return [
'http://www.pass-e.com/proxy/index.php?page=%(n)01d' % {'n': i}
for i in range(1, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_4 refactored with the following changes:

Comment on lines -476 to +461
ret = []
for i in range(1, page):
ret.append('http://www.ipfree.cn/index2.asp?page=%(num)01d' % {'num': i})
return ret
return [
'http://www.ipfree.cn/index2.asp?page=%(num)01d' % {'num': i}
for i in range(1, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_5 refactored with the following changes:

Comment on lines -508 to +493
ret = []
for i in range(1, page):
ret.append('http://www.cnproxy.com/proxy%(num)01d.html' % {'num': i})
return ret
return [
'http://www.cnproxy.com/proxy%(num)01d.html' % {'num': i}
for i in range(1, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_6 refactored with the following changes:

Comment on lines +507 to -528
type = -1 # 该网站未提供代理服务器类型
for match in matches:
ip = match[0]
port = match[1]
type = -1 # 该网站未提供代理服务器类型

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function parse_page_6 refactored with the following changes:

Comment on lines -595 to +580
ret = []
for i in range(0, page):
ret.append('http://proxylist.sakura.ne.jp/index.htm?pages=%(n)01d' % {'n': i})
return ret
return [
'http://proxylist.sakura.ne.jp/index.htm?pages=%(n)01d' % {'n': i}
for i in range(0, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_9 refactored with the following changes:

Comment on lines -616 to +598
if (type == 'Anonymous'):
type = 1
else:
type = -1
type = 1 if (type == 'Anonymous') else -1

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function parse_page_9 refactored with the following changes:

Comment on lines -634 to +616
ret = []
for i in range(1, page):
ret.append('http://www.publicproxyservers.com/page%(n)01d.html' % {'n': i})
return ret
return [
'http://www.publicproxyservers.com/page%(n)01d.html' % {'n': i}
for i in range(1, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_10 refactored with the following changes:

Comment on lines -678 to +667
ret = []
for i in range(1, page):
ret.append('http://www.my-proxy.com/list/proxy.php?list=%(n)01d' % {'n': i})

ret.append('http://www.my-proxy.com/list/proxy.php?list=s1')
ret.append('http://www.my-proxy.com/list/proxy.php?list=s2')
ret.append('http://www.my-proxy.com/list/proxy.php?list=s3')
ret = [
'http://www.my-proxy.com/list/proxy.php?list=%(n)01d' % {'n': i}
for i in range(1, page)
]
ret.extend(
(
'http://www.my-proxy.com/list/proxy.php?list=s1',
'http://www.my-proxy.com/list/proxy.php?list=s2',
'http://www.my-proxy.com/list/proxy.php?list=s3',
)
)

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_11 refactored with the following changes:

urls = []
for page in range(beg, end):
urls.append('http://www.baidu.com?&page=%d' % page)
urls = ['http://www.baidu.com?&page=%d' % page for page in range(beg, end)]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function run_spider refactored with the following changes:

Comment on lines -92 to +90
res = []
for i in range(len(s) - 1):
res.append((s[i], s[i + 1]))
res = [(s[i], s[i + 1]) for i in range(len(s) - 1)]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function main refactored with the following changes:

Comment on lines -64 to -68
proxies = {
'http': '{proxy_type}://{ip}:{port}'.format(proxy_type=proxy_type, ip=ip, port=port),
'https': '{proxy_type}://{ip}:{port}'.format(proxy_type=proxy_type, ip=ip, port=port),
return {
'http': '{proxy_type}://{ip}:{port}'.format(
proxy_type=proxy_type, ip=ip, port=port
),
'https': '{proxy_type}://{ip}:{port}'.format(
proxy_type=proxy_type, ip=ip, port=port
),
}
return proxies

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function get_proxy_dict refactored with the following changes:

urls = []
for page in range(1, 1000):
urls.append('http://www.jb51.net/article/%s.htm' % page)
urls = [f'http://www.jb51.net/article/{page}.htm' for page in range(1, 1000)]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function main refactored with the following changes:

Comment threadcrawler/src/test.py
class TestSyncSpider(SyncSpider):
def handle_html(self, url, html):
print(html)
pass

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function TestSyncSpider.handle_html refactored with the following changes:

Comment threadcrawler/src/test.py
class TestAsyncSpider(AsyncSpider):
def handle_html(self, url, html):
print(html)
pass

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function TestAsyncSpider.handle_html refactored with the following changes:

Comment threadcrawler/src/test.py
Comment on lines -24 to +25
urls = []
for page in range(1, 1000):
#urls.append('http://www.jb51.net/article/%s.htm' % page)
urls.append('http://www.imooc.com/data/check_f.php?page=%d'%page)
urls = [
'http://www.imooc.com/data/check_f.php?page=%d' % page
for page in range(1, 1000)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Lines 24-27 refactored with the following changes:

This removes the following comments ( why? ):

#urls.append('http://www.jb51.net/article/%s.htm' % page)

Comment on lines -64 to +69
func_name = 'get_' + field
func_name = f'get_{field}'
xpath_str = self.xpath_dict.get(field)
if hasattr(self, func_name):
return getattr(self, func_name)(xpath_str)
else:
self.logger.debug(field, self.url)
return self.parser.xpath(xpath_str)[0].strip() if xpath_str else ''
self.logger.debug(field, self.url)
return self.parser.xpath(xpath_str)[0].strip() if xpath_str else ''

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function XpathCrawler.get_field refactored with the following changes:

Comment on lines -73 to +75
xpath_result = {}
for field, xpath_string in self.xpath_dict.items():
xpath_result[field] = makes(self.get_field(field)) # to utf8
return xpath_result
return {
field: makes(self.get_field(field))
for field, xpath_string in self.xpath_dict.items()
}

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function XpathCrawler.get_result refactored with the following changes:

This removes the following comments ( why? ):

# to utf8

Comment on lines -93 to +92
category_urls = []
for href in category_hrefs:
category_urls.append(urljoin(self.domain, href))
category_urls = [urljoin(self.domain, href) for href in category_hrefs]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function FantasyhairbuySite.generate_category_urls refactored with the following changes:

Comment on lines -11 to +14
backup_name = os.path.basename(file_path) + '_' + datetime.datetime.now().strftime('%Y%m%d%H%M%S')
backup_name = (
f'{os.path.basename(file_path)}_'
+ datetime.datetime.now().strftime('%Y%m%d%H%M%S')
)

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function backup_file refactored with the following changes:

Comment on lines -44 to +47
backup_list = glob.glob(os.path.join(path, pattern + '_*'))
backup_list = glob.glob(os.path.join(path, f'{pattern}_*'))

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function main refactored with the following changes:


# if it's anything else, return it in its original form
return data
return data.encode('utf-8') if str(type(data)) == "<type 'unicode'>" else data

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function _byteify refactored with the following changes:

This removes the following comments ( why? ):

# if it's anything else, return it in its original form

try:
res = query.find()
return res
return query.find()

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function LeanCloudApi.get_skip_obj_list refactored with the following changes:

Comment on lines -49 to +48
img_info_url = img_url + '?imageInfo'
img_info_url = f'{img_url}?imageInfo'

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function LeanCloudApi.add_img_info refactored with the following changes:

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Dev (Sourcery refactored) - #18

Open
sourcery-ai[bot] wants to merge 1 commit into
devfrom
sourcery/dev
Open

Dev (Sourcery refactored)#18
sourcery-ai[bot] wants to merge 1 commit into
devfrom
sourcery/dev

Conversation

@sourcery-ai

Copy link
Copy Markdown

Pull Request #17 refactored by Sourcery.

If you're happy with these changes, merge this Pull Request using the Squash and merge strategy.

NOTE: As code is pushed to the original Pull Request, Sourcery will
re-run and update (force-push) this Pull Request with new refactorings as
necessary. If Sourcery finds no refactorings at any point, this Pull Request
will be closed automatically.

See our documentation here.

Run Sourcery locally

Reduce the feedback loop during development by using the Sourcery editor plugin:

Review changes via command line

To manually merge these changes, make sure you're on the dev branch, then run:

git fetch origin sourcery/dev
git merge --ff-only FETCH_HEAD
git reset HEAD^

Help us improve this pull request!

@sourcery-aisourcery-aiBot left a comment

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Due to GitHub API limits, only the first 60 comments can be shown.

Comment threadSchedule/demo4.py

def greet(name):
print('Hello {}'.format(name))
print(f'Hello {name}')

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function greet refactored with the following changes:

Comment threadSchedule/demo8.py
"""
def job1():
print("I'm running on threads %s" % threading.current_thread())
print(f"I'm running on threads {threading.current_thread()}")

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function job1 refactored with the following changes:

Comment threadSchedule/demo8.py
print(f"I'm running on threads {threading.current_thread()}")
def job2():
print("I'm running on threads %s" % threading.current_thread())
print(f"I'm running on threads {threading.current_thread()}")

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function job2 refactored with the following changes:

Comment threadSchedule/demo8.py
print(f"I'm running on threads {threading.current_thread()}")
def job3():
print("I'm running on threads %s" % threading.current_thread())
print(f"I'm running on threads {threading.current_thread()}")

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function job3 refactored with the following changes:

Comment on lines -13 to +15
CHROME_DRIVER_MAPPING_FILE = r"{}\mapping.json".format(CHROME_DRIVER_FOLDER)
CHROME_DRIVER_EXE = r"{}\chromedriver.exe".format(CHROME_DRIVER_FOLDER)
CHROME_DRIVER_ZIP = r"{}\chromedriver_win32.zip".format(CHROME_DRIVER_FOLDER)
CHROME_DRIVER_MAPPING_FILE = f"{CHROME_DRIVER_FOLDER}\mapping.json"
CHROME_DRIVER_EXE = f"{CHROME_DRIVER_FOLDER}\chromedriver.exe"
CHROME_DRIVER_ZIP = f"{CHROME_DRIVER_FOLDER}\chromedriver_win32.zip"

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Lines 13-15 refactored with the following changes:

Comment on lines -386 to +371
ret = []
for i in range(1, page):
ret.append('http://www.samair.ru/proxy/proxy-%(num)02d.htm' % {'num': i})
return ret
return [
'http://www.samair.ru/proxy/proxy-%(num)02d.htm' % {'num': i}
for i in range(1, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_3 refactored with the following changes:

Comment on lines -406 to +388
ip = match[0] + "." + match[1] + match[2]
ip = f"{match[0]}.{match[1]}{match[2]}"

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function parse_page_3 refactored with the following changes:

Comment on lines -432 to +417
ret = []
for i in range(1, page):
ret.append('http://www.pass-e.com/proxy/index.php?page=%(n)01d' % {'n': i})
return ret
return [
'http://www.pass-e.com/proxy/index.php?page=%(n)01d' % {'n': i}
for i in range(1, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_4 refactored with the following changes:

Comment on lines -476 to +461
ret = []
for i in range(1, page):
ret.append('http://www.ipfree.cn/index2.asp?page=%(num)01d' % {'num': i})
return ret
return [
'http://www.ipfree.cn/index2.asp?page=%(num)01d' % {'num': i}
for i in range(1, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_5 refactored with the following changes:

Comment on lines -508 to +493
ret = []
for i in range(1, page):
ret.append('http://www.cnproxy.com/proxy%(num)01d.html' % {'num': i})
return ret
return [
'http://www.cnproxy.com/proxy%(num)01d.html' % {'num': i}
for i in range(1, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_6 refactored with the following changes:

Comment on lines +507 to -528
type = -1 # 该网站未提供代理服务器类型
for match in matches:
ip = match[0]
port = match[1]
type = -1 # 该网站未提供代理服务器类型

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function parse_page_6 refactored with the following changes:

Comment on lines -595 to +580
ret = []
for i in range(0, page):
ret.append('http://proxylist.sakura.ne.jp/index.htm?pages=%(n)01d' % {'n': i})
return ret
return [
'http://proxylist.sakura.ne.jp/index.htm?pages=%(n)01d' % {'n': i}
for i in range(0, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_9 refactored with the following changes:

Comment on lines -616 to +598
if (type == 'Anonymous'):
type = 1
else:
type = -1
type = 1 if (type == 'Anonymous') else -1

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function parse_page_9 refactored with the following changes:

Comment on lines -634 to +616
ret = []
for i in range(1, page):
ret.append('http://www.publicproxyservers.com/page%(n)01d.html' % {'n': i})
return ret
return [
'http://www.publicproxyservers.com/page%(n)01d.html' % {'n': i}
for i in range(1, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_10 refactored with the following changes:

Comment on lines -678 to +667
ret = []
for i in range(1, page):
ret.append('http://www.my-proxy.com/list/proxy.php?list=%(n)01d' % {'n': i})

ret.append('http://www.my-proxy.com/list/proxy.php?list=s1')
ret.append('http://www.my-proxy.com/list/proxy.php?list=s2')
ret.append('http://www.my-proxy.com/list/proxy.php?list=s3')
ret = [
'http://www.my-proxy.com/list/proxy.php?list=%(n)01d' % {'n': i}
for i in range(1, page)
]
ret.extend(
(
'http://www.my-proxy.com/list/proxy.php?list=s1',
'http://www.my-proxy.com/list/proxy.php?list=s2',
'http://www.my-proxy.com/list/proxy.php?list=s3',
)
)

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_11 refactored with the following changes:

urls = []
for page in range(beg, end):
urls.append('http://www.baidu.com?&page=%d' % page)
urls = ['http://www.baidu.com?&page=%d' % page for page in range(beg, end)]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function run_spider refactored with the following changes:

Comment on lines -92 to +90
res = []
for i in range(len(s) - 1):
res.append((s[i], s[i + 1]))
res = [(s[i], s[i + 1]) for i in range(len(s) - 1)]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function main refactored with the following changes:

Comment on lines -64 to -68
proxies = {
'http': '{proxy_type}://{ip}:{port}'.format(proxy_type=proxy_type, ip=ip, port=port),
'https': '{proxy_type}://{ip}:{port}'.format(proxy_type=proxy_type, ip=ip, port=port),
return {
'http': '{proxy_type}://{ip}:{port}'.format(
proxy_type=proxy_type, ip=ip, port=port
),
'https': '{proxy_type}://{ip}:{port}'.format(
proxy_type=proxy_type, ip=ip, port=port
),
}
return proxies

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function get_proxy_dict refactored with the following changes:

urls = []
for page in range(1, 1000):
urls.append('http://www.jb51.net/article/%s.htm' % page)
urls = [f'http://www.jb51.net/article/{page}.htm' for page in range(1, 1000)]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function main refactored with the following changes:

Comment threadcrawler/src/test.py
class TestSyncSpider(SyncSpider):
def handle_html(self, url, html):
print(html)
pass

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function TestSyncSpider.handle_html refactored with the following changes:

Comment threadcrawler/src/test.py
class TestAsyncSpider(AsyncSpider):
def handle_html(self, url, html):
print(html)
pass

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function TestAsyncSpider.handle_html refactored with the following changes:

Comment threadcrawler/src/test.py
Comment on lines -24 to +25
urls = []
for page in range(1, 1000):
#urls.append('http://www.jb51.net/article/%s.htm' % page)
urls.append('http://www.imooc.com/data/check_f.php?page=%d'%page)
urls = [
'http://www.imooc.com/data/check_f.php?page=%d' % page
for page in range(1, 1000)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Lines 24-27 refactored with the following changes:

This removes the following comments ( why? ):

#urls.append('http://www.jb51.net/article/%s.htm' % page)

Comment on lines -64 to +69
func_name = 'get_' + field
func_name = f'get_{field}'
xpath_str = self.xpath_dict.get(field)
if hasattr(self, func_name):
return getattr(self, func_name)(xpath_str)
else:
self.logger.debug(field, self.url)
return self.parser.xpath(xpath_str)[0].strip() if xpath_str else ''
self.logger.debug(field, self.url)
return self.parser.xpath(xpath_str)[0].strip() if xpath_str else ''

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function XpathCrawler.get_field refactored with the following changes:

Comment on lines -73 to +75
xpath_result = {}
for field, xpath_string in self.xpath_dict.items():
xpath_result[field] = makes(self.get_field(field)) # to utf8
return xpath_result
return {
field: makes(self.get_field(field))
for field, xpath_string in self.xpath_dict.items()
}

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function XpathCrawler.get_result refactored with the following changes:

This removes the following comments ( why? ):

# to utf8

Comment on lines -93 to +92
category_urls = []
for href in category_hrefs:
category_urls.append(urljoin(self.domain, href))
category_urls = [urljoin(self.domain, href) for href in category_hrefs]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function FantasyhairbuySite.generate_category_urls refactored with the following changes:

Comment on lines -11 to +14
backup_name = os.path.basename(file_path) + '_' + datetime.datetime.now().strftime('%Y%m%d%H%M%S')
backup_name = (
f'{os.path.basename(file_path)}_'
+ datetime.datetime.now().strftime('%Y%m%d%H%M%S')
)

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function backup_file refactored with the following changes:

Comment on lines -44 to +47
backup_list = glob.glob(os.path.join(path, pattern + '_*'))
backup_list = glob.glob(os.path.join(path, f'{pattern}_*'))

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function main refactored with the following changes:


# if it's anything else, return it in its original form
return data
return data.encode('utf-8') if str(type(data)) == "<type 'unicode'>" else data

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function _byteify refactored with the following changes:

This removes the following comments ( why? ):

# if it's anything else, return it in its original form

try:
res = query.find()
return res
return query.find()

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function LeanCloudApi.get_skip_obj_list refactored with the following changes:

Comment on lines -49 to +48
img_info_url = img_url + '?imageInfo'
img_info_url = f'{img_url}?imageInfo'

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function LeanCloudApi.add_img_info refactored with the following changes:

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Dev (Sourcery refactored) - #18

Open
sourcery-ai[bot] wants to merge 1 commit into
devfrom
sourcery/dev
Open

Dev (Sourcery refactored)#18
sourcery-ai[bot] wants to merge 1 commit into
devfrom
sourcery/dev

Conversation

@sourcery-ai

Copy link
Copy Markdown

Pull Request #17 refactored by Sourcery.

If you're happy with these changes, merge this Pull Request using the Squash and merge strategy.

NOTE: As code is pushed to the original Pull Request, Sourcery will
re-run and update (force-push) this Pull Request with new refactorings as
necessary. If Sourcery finds no refactorings at any point, this Pull Request
will be closed automatically.

See our documentation here.

Run Sourcery locally

Reduce the feedback loop during development by using the Sourcery editor plugin:

Review changes via command line

To manually merge these changes, make sure you're on the dev branch, then run:

git fetch origin sourcery/dev
git merge --ff-only FETCH_HEAD
git reset HEAD^

Help us improve this pull request!

@sourcery-aisourcery-aiBot left a comment

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Due to GitHub API limits, only the first 60 comments can be shown.

Comment threadSchedule/demo4.py

def greet(name):
print('Hello {}'.format(name))
print(f'Hello {name}')

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function greet refactored with the following changes:

Comment threadSchedule/demo8.py
"""
def job1():
print("I'm running on threads %s" % threading.current_thread())
print(f"I'm running on threads {threading.current_thread()}")

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function job1 refactored with the following changes:

Comment threadSchedule/demo8.py
print(f"I'm running on threads {threading.current_thread()}")
def job2():
print("I'm running on threads %s" % threading.current_thread())
print(f"I'm running on threads {threading.current_thread()}")

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function job2 refactored with the following changes:

Comment threadSchedule/demo8.py
print(f"I'm running on threads {threading.current_thread()}")
def job3():
print("I'm running on threads %s" % threading.current_thread())
print(f"I'm running on threads {threading.current_thread()}")

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function job3 refactored with the following changes:

Comment on lines -13 to +15
CHROME_DRIVER_MAPPING_FILE = r"{}\mapping.json".format(CHROME_DRIVER_FOLDER)
CHROME_DRIVER_EXE = r"{}\chromedriver.exe".format(CHROME_DRIVER_FOLDER)
CHROME_DRIVER_ZIP = r"{}\chromedriver_win32.zip".format(CHROME_DRIVER_FOLDER)
CHROME_DRIVER_MAPPING_FILE = f"{CHROME_DRIVER_FOLDER}\mapping.json"
CHROME_DRIVER_EXE = f"{CHROME_DRIVER_FOLDER}\chromedriver.exe"
CHROME_DRIVER_ZIP = f"{CHROME_DRIVER_FOLDER}\chromedriver_win32.zip"

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Lines 13-15 refactored with the following changes:

Comment on lines -386 to +371
ret = []
for i in range(1, page):
ret.append('http://www.samair.ru/proxy/proxy-%(num)02d.htm' % {'num': i})
return ret
return [
'http://www.samair.ru/proxy/proxy-%(num)02d.htm' % {'num': i}
for i in range(1, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_3 refactored with the following changes:

Comment on lines -406 to +388
ip = match[0] + "." + match[1] + match[2]
ip = f"{match[0]}.{match[1]}{match[2]}"

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function parse_page_3 refactored with the following changes:

Comment on lines -432 to +417
ret = []
for i in range(1, page):
ret.append('http://www.pass-e.com/proxy/index.php?page=%(n)01d' % {'n': i})
return ret
return [
'http://www.pass-e.com/proxy/index.php?page=%(n)01d' % {'n': i}
for i in range(1, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_4 refactored with the following changes:

Comment on lines -476 to +461
ret = []
for i in range(1, page):
ret.append('http://www.ipfree.cn/index2.asp?page=%(num)01d' % {'num': i})
return ret
return [
'http://www.ipfree.cn/index2.asp?page=%(num)01d' % {'num': i}
for i in range(1, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_5 refactored with the following changes:

Comment on lines -508 to +493
ret = []
for i in range(1, page):
ret.append('http://www.cnproxy.com/proxy%(num)01d.html' % {'num': i})
return ret
return [
'http://www.cnproxy.com/proxy%(num)01d.html' % {'num': i}
for i in range(1, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_6 refactored with the following changes:

Comment on lines +507 to -528
type = -1 # 该网站未提供代理服务器类型
for match in matches:
ip = match[0]
port = match[1]
type = -1 # 该网站未提供代理服务器类型

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function parse_page_6 refactored with the following changes:

Comment on lines -595 to +580
ret = []
for i in range(0, page):
ret.append('http://proxylist.sakura.ne.jp/index.htm?pages=%(n)01d' % {'n': i})
return ret
return [
'http://proxylist.sakura.ne.jp/index.htm?pages=%(n)01d' % {'n': i}
for i in range(0, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_9 refactored with the following changes:

Comment on lines -616 to +598
if (type == 'Anonymous'):
type = 1
else:
type = -1
type = 1 if (type == 'Anonymous') else -1

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function parse_page_9 refactored with the following changes:

Comment on lines -634 to +616
ret = []
for i in range(1, page):
ret.append('http://www.publicproxyservers.com/page%(n)01d.html' % {'n': i})
return ret
return [
'http://www.publicproxyservers.com/page%(n)01d.html' % {'n': i}
for i in range(1, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_10 refactored with the following changes:

Comment on lines -678 to +667
ret = []
for i in range(1, page):
ret.append('http://www.my-proxy.com/list/proxy.php?list=%(n)01d' % {'n': i})

ret.append('http://www.my-proxy.com/list/proxy.php?list=s1')
ret.append('http://www.my-proxy.com/list/proxy.php?list=s2')
ret.append('http://www.my-proxy.com/list/proxy.php?list=s3')
ret = [
'http://www.my-proxy.com/list/proxy.php?list=%(n)01d' % {'n': i}
for i in range(1, page)
]
ret.extend(
(
'http://www.my-proxy.com/list/proxy.php?list=s1',
'http://www.my-proxy.com/list/proxy.php?list=s2',
'http://www.my-proxy.com/list/proxy.php?list=s3',
)
)

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_11 refactored with the following changes:

urls = []
for page in range(beg, end):
urls.append('http://www.baidu.com?&page=%d' % page)
urls = ['http://www.baidu.com?&page=%d' % page for page in range(beg, end)]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function run_spider refactored with the following changes:

Comment on lines -92 to +90
res = []
for i in range(len(s) - 1):
res.append((s[i], s[i + 1]))
res = [(s[i], s[i + 1]) for i in range(len(s) - 1)]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function main refactored with the following changes:

Comment on lines -64 to -68
proxies = {
'http': '{proxy_type}://{ip}:{port}'.format(proxy_type=proxy_type, ip=ip, port=port),
'https': '{proxy_type}://{ip}:{port}'.format(proxy_type=proxy_type, ip=ip, port=port),
return {
'http': '{proxy_type}://{ip}:{port}'.format(
proxy_type=proxy_type, ip=ip, port=port
),
'https': '{proxy_type}://{ip}:{port}'.format(
proxy_type=proxy_type, ip=ip, port=port
),
}
return proxies

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function get_proxy_dict refactored with the following changes:

urls = []
for page in range(1, 1000):
urls.append('http://www.jb51.net/article/%s.htm' % page)
urls = [f'http://www.jb51.net/article/{page}.htm' for page in range(1, 1000)]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function main refactored with the following changes:

Comment threadcrawler/src/test.py
class TestSyncSpider(SyncSpider):
def handle_html(self, url, html):
print(html)
pass

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function TestSyncSpider.handle_html refactored with the following changes:

Comment threadcrawler/src/test.py
class TestAsyncSpider(AsyncSpider):
def handle_html(self, url, html):
print(html)
pass

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function TestAsyncSpider.handle_html refactored with the following changes:

Comment threadcrawler/src/test.py
Comment on lines -24 to +25
urls = []
for page in range(1, 1000):
#urls.append('http://www.jb51.net/article/%s.htm' % page)
urls.append('http://www.imooc.com/data/check_f.php?page=%d'%page)
urls = [
'http://www.imooc.com/data/check_f.php?page=%d' % page
for page in range(1, 1000)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Lines 24-27 refactored with the following changes:

This removes the following comments ( why? ):

#urls.append('http://www.jb51.net/article/%s.htm' % page)

Comment on lines -64 to +69
func_name = 'get_' + field
func_name = f'get_{field}'
xpath_str = self.xpath_dict.get(field)
if hasattr(self, func_name):
return getattr(self, func_name)(xpath_str)
else:
self.logger.debug(field, self.url)
return self.parser.xpath(xpath_str)[0].strip() if xpath_str else ''
self.logger.debug(field, self.url)
return self.parser.xpath(xpath_str)[0].strip() if xpath_str else ''

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function XpathCrawler.get_field refactored with the following changes:

Comment on lines -73 to +75
xpath_result = {}
for field, xpath_string in self.xpath_dict.items():
xpath_result[field] = makes(self.get_field(field)) # to utf8
return xpath_result
return {
field: makes(self.get_field(field))
for field, xpath_string in self.xpath_dict.items()
}

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function XpathCrawler.get_result refactored with the following changes:

This removes the following comments ( why? ):

# to utf8

Comment on lines -93 to +92
category_urls = []
for href in category_hrefs:
category_urls.append(urljoin(self.domain, href))
category_urls = [urljoin(self.domain, href) for href in category_hrefs]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function FantasyhairbuySite.generate_category_urls refactored with the following changes:

Comment on lines -11 to +14
backup_name = os.path.basename(file_path) + '_' + datetime.datetime.now().strftime('%Y%m%d%H%M%S')
backup_name = (
f'{os.path.basename(file_path)}_'
+ datetime.datetime.now().strftime('%Y%m%d%H%M%S')
)

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function backup_file refactored with the following changes:

Comment on lines -44 to +47
backup_list = glob.glob(os.path.join(path, pattern + '_*'))
backup_list = glob.glob(os.path.join(path, f'{pattern}_*'))

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function main refactored with the following changes:


# if it's anything else, return it in its original form
return data
return data.encode('utf-8') if str(type(data)) == "<type 'unicode'>" else data

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function _byteify refactored with the following changes:

This removes the following comments ( why? ):

# if it's anything else, return it in its original form

try:
res = query.find()
return res
return query.find()

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function LeanCloudApi.get_skip_obj_list refactored with the following changes:

Comment on lines -49 to +48
img_info_url = img_url + '?imageInfo'
img_info_url = f'{img_url}?imageInfo'

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function LeanCloudApi.add_img_info refactored with the following changes:

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Dev (Sourcery refactored) - #18

Open
sourcery-ai[bot] wants to merge 1 commit into
devfrom
sourcery/dev
Open

Dev (Sourcery refactored)#18
sourcery-ai[bot] wants to merge 1 commit into
devfrom
sourcery/dev

Conversation

@sourcery-ai

Copy link
Copy Markdown

Pull Request #17 refactored by Sourcery.

If you're happy with these changes, merge this Pull Request using the Squash and merge strategy.

NOTE: As code is pushed to the original Pull Request, Sourcery will
re-run and update (force-push) this Pull Request with new refactorings as
necessary. If Sourcery finds no refactorings at any point, this Pull Request
will be closed automatically.

See our documentation here.

Run Sourcery locally

Reduce the feedback loop during development by using the Sourcery editor plugin:

Review changes via command line

To manually merge these changes, make sure you're on the dev branch, then run:

git fetch origin sourcery/dev
git merge --ff-only FETCH_HEAD
git reset HEAD^

Help us improve this pull request!

@sourcery-aisourcery-aiBot left a comment

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Due to GitHub API limits, only the first 60 comments can be shown.

Comment threadSchedule/demo4.py

def greet(name):
print('Hello {}'.format(name))
print(f'Hello {name}')

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function greet refactored with the following changes:

Comment threadSchedule/demo8.py
"""
def job1():
print("I'm running on threads %s" % threading.current_thread())
print(f"I'm running on threads {threading.current_thread()}")

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function job1 refactored with the following changes:

Comment threadSchedule/demo8.py
print(f"I'm running on threads {threading.current_thread()}")
def job2():
print("I'm running on threads %s" % threading.current_thread())
print(f"I'm running on threads {threading.current_thread()}")

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function job2 refactored with the following changes:

Comment threadSchedule/demo8.py
print(f"I'm running on threads {threading.current_thread()}")
def job3():
print("I'm running on threads %s" % threading.current_thread())
print(f"I'm running on threads {threading.current_thread()}")

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function job3 refactored with the following changes:

Comment on lines -13 to +15
CHROME_DRIVER_MAPPING_FILE = r"{}\mapping.json".format(CHROME_DRIVER_FOLDER)
CHROME_DRIVER_EXE = r"{}\chromedriver.exe".format(CHROME_DRIVER_FOLDER)
CHROME_DRIVER_ZIP = r"{}\chromedriver_win32.zip".format(CHROME_DRIVER_FOLDER)
CHROME_DRIVER_MAPPING_FILE = f"{CHROME_DRIVER_FOLDER}\mapping.json"
CHROME_DRIVER_EXE = f"{CHROME_DRIVER_FOLDER}\chromedriver.exe"
CHROME_DRIVER_ZIP = f"{CHROME_DRIVER_FOLDER}\chromedriver_win32.zip"

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Lines 13-15 refactored with the following changes:

Comment on lines -386 to +371
ret = []
for i in range(1, page):
ret.append('http://www.samair.ru/proxy/proxy-%(num)02d.htm' % {'num': i})
return ret
return [
'http://www.samair.ru/proxy/proxy-%(num)02d.htm' % {'num': i}
for i in range(1, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_3 refactored with the following changes:

Comment on lines -406 to +388
ip = match[0] + "." + match[1] + match[2]
ip = f"{match[0]}.{match[1]}{match[2]}"

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function parse_page_3 refactored with the following changes:

Comment on lines -432 to +417
ret = []
for i in range(1, page):
ret.append('http://www.pass-e.com/proxy/index.php?page=%(n)01d' % {'n': i})
return ret
return [
'http://www.pass-e.com/proxy/index.php?page=%(n)01d' % {'n': i}
for i in range(1, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_4 refactored with the following changes:

Comment on lines -476 to +461
ret = []
for i in range(1, page):
ret.append('http://www.ipfree.cn/index2.asp?page=%(num)01d' % {'num': i})
return ret
return [
'http://www.ipfree.cn/index2.asp?page=%(num)01d' % {'num': i}
for i in range(1, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_5 refactored with the following changes:

Comment on lines -508 to +493
ret = []
for i in range(1, page):
ret.append('http://www.cnproxy.com/proxy%(num)01d.html' % {'num': i})
return ret
return [
'http://www.cnproxy.com/proxy%(num)01d.html' % {'num': i}
for i in range(1, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_6 refactored with the following changes:

Comment on lines +507 to -528
type = -1 # 该网站未提供代理服务器类型
for match in matches:
ip = match[0]
port = match[1]
type = -1 # 该网站未提供代理服务器类型

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function parse_page_6 refactored with the following changes:

Comment on lines -595 to +580
ret = []
for i in range(0, page):
ret.append('http://proxylist.sakura.ne.jp/index.htm?pages=%(n)01d' % {'n': i})
return ret
return [
'http://proxylist.sakura.ne.jp/index.htm?pages=%(n)01d' % {'n': i}
for i in range(0, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_9 refactored with the following changes:

Comment on lines -616 to +598
if (type == 'Anonymous'):
type = 1
else:
type = -1
type = 1 if (type == 'Anonymous') else -1

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function parse_page_9 refactored with the following changes:

Comment on lines -634 to +616
ret = []
for i in range(1, page):
ret.append('http://www.publicproxyservers.com/page%(n)01d.html' % {'n': i})
return ret
return [
'http://www.publicproxyservers.com/page%(n)01d.html' % {'n': i}
for i in range(1, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_10 refactored with the following changes:

Comment on lines -678 to +667
ret = []
for i in range(1, page):
ret.append('http://www.my-proxy.com/list/proxy.php?list=%(n)01d' % {'n': i})

ret.append('http://www.my-proxy.com/list/proxy.php?list=s1')
ret.append('http://www.my-proxy.com/list/proxy.php?list=s2')
ret.append('http://www.my-proxy.com/list/proxy.php?list=s3')
ret = [
'http://www.my-proxy.com/list/proxy.php?list=%(n)01d' % {'n': i}
for i in range(1, page)
]
ret.extend(
(
'http://www.my-proxy.com/list/proxy.php?list=s1',
'http://www.my-proxy.com/list/proxy.php?list=s2',
'http://www.my-proxy.com/list/proxy.php?list=s3',
)
)

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_11 refactored with the following changes:

urls = []
for page in range(beg, end):
urls.append('http://www.baidu.com?&page=%d' % page)
urls = ['http://www.baidu.com?&page=%d' % page for page in range(beg, end)]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function run_spider refactored with the following changes:

Comment on lines -92 to +90
res = []
for i in range(len(s) - 1):
res.append((s[i], s[i + 1]))
res = [(s[i], s[i + 1]) for i in range(len(s) - 1)]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function main refactored with the following changes:

Comment on lines -64 to -68
proxies = {
'http': '{proxy_type}://{ip}:{port}'.format(proxy_type=proxy_type, ip=ip, port=port),
'https': '{proxy_type}://{ip}:{port}'.format(proxy_type=proxy_type, ip=ip, port=port),
return {
'http': '{proxy_type}://{ip}:{port}'.format(
proxy_type=proxy_type, ip=ip, port=port
),
'https': '{proxy_type}://{ip}:{port}'.format(
proxy_type=proxy_type, ip=ip, port=port
),
}
return proxies

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function get_proxy_dict refactored with the following changes:

urls = []
for page in range(1, 1000):
urls.append('http://www.jb51.net/article/%s.htm' % page)
urls = [f'http://www.jb51.net/article/{page}.htm' for page in range(1, 1000)]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function main refactored with the following changes:

Comment threadcrawler/src/test.py
class TestSyncSpider(SyncSpider):
def handle_html(self, url, html):
print(html)
pass

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function TestSyncSpider.handle_html refactored with the following changes:

Comment threadcrawler/src/test.py
class TestAsyncSpider(AsyncSpider):
def handle_html(self, url, html):
print(html)
pass

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function TestAsyncSpider.handle_html refactored with the following changes:

Comment threadcrawler/src/test.py
Comment on lines -24 to +25
urls = []
for page in range(1, 1000):
#urls.append('http://www.jb51.net/article/%s.htm' % page)
urls.append('http://www.imooc.com/data/check_f.php?page=%d'%page)
urls = [
'http://www.imooc.com/data/check_f.php?page=%d' % page
for page in range(1, 1000)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Lines 24-27 refactored with the following changes:

This removes the following comments ( why? ):

#urls.append('http://www.jb51.net/article/%s.htm' % page)

Comment on lines -64 to +69
func_name = 'get_' + field
func_name = f'get_{field}'
xpath_str = self.xpath_dict.get(field)
if hasattr(self, func_name):
return getattr(self, func_name)(xpath_str)
else:
self.logger.debug(field, self.url)
return self.parser.xpath(xpath_str)[0].strip() if xpath_str else ''
self.logger.debug(field, self.url)
return self.parser.xpath(xpath_str)[0].strip() if xpath_str else ''

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function XpathCrawler.get_field refactored with the following changes:

Comment on lines -73 to +75
xpath_result = {}
for field, xpath_string in self.xpath_dict.items():
xpath_result[field] = makes(self.get_field(field)) # to utf8
return xpath_result
return {
field: makes(self.get_field(field))
for field, xpath_string in self.xpath_dict.items()
}

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function XpathCrawler.get_result refactored with the following changes:

This removes the following comments ( why? ):

# to utf8

Comment on lines -93 to +92
category_urls = []
for href in category_hrefs:
category_urls.append(urljoin(self.domain, href))
category_urls = [urljoin(self.domain, href) for href in category_hrefs]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function FantasyhairbuySite.generate_category_urls refactored with the following changes:

Comment on lines -11 to +14
backup_name = os.path.basename(file_path) + '_' + datetime.datetime.now().strftime('%Y%m%d%H%M%S')
backup_name = (
f'{os.path.basename(file_path)}_'
+ datetime.datetime.now().strftime('%Y%m%d%H%M%S')
)

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function backup_file refactored with the following changes:

Comment on lines -44 to +47
backup_list = glob.glob(os.path.join(path, pattern + '_*'))
backup_list = glob.glob(os.path.join(path, f'{pattern}_*'))

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function main refactored with the following changes:


# if it's anything else, return it in its original form
return data
return data.encode('utf-8') if str(type(data)) == "<type 'unicode'>" else data

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function _byteify refactored with the following changes:

This removes the following comments ( why? ):

# if it's anything else, return it in its original form

try:
res = query.find()
return res
return query.find()

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function LeanCloudApi.get_skip_obj_list refactored with the following changes:

Comment on lines -49 to +48
img_info_url = img_url + '?imageInfo'
img_info_url = f'{img_url}?imageInfo'

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function LeanCloudApi.add_img_info refactored with the following changes:

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Dev (Sourcery refactored) - #18

Open
sourcery-ai[bot] wants to merge 1 commit into
devfrom
sourcery/dev
Open

Dev (Sourcery refactored)#18
sourcery-ai[bot] wants to merge 1 commit into
devfrom
sourcery/dev

Conversation

@sourcery-ai

Copy link
Copy Markdown

Pull Request #17 refactored by Sourcery.

If you're happy with these changes, merge this Pull Request using the Squash and merge strategy.

NOTE: As code is pushed to the original Pull Request, Sourcery will
re-run and update (force-push) this Pull Request with new refactorings as
necessary. If Sourcery finds no refactorings at any point, this Pull Request
will be closed automatically.

See our documentation here.

Run Sourcery locally

Reduce the feedback loop during development by using the Sourcery editor plugin:

Review changes via command line

To manually merge these changes, make sure you're on the dev branch, then run:

git fetch origin sourcery/dev
git merge --ff-only FETCH_HEAD
git reset HEAD^

Help us improve this pull request!

@sourcery-aisourcery-aiBot left a comment

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Due to GitHub API limits, only the first 60 comments can be shown.

Comment threadSchedule/demo4.py

def greet(name):
print('Hello {}'.format(name))
print(f'Hello {name}')

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function greet refactored with the following changes:

Comment threadSchedule/demo8.py
"""
def job1():
print("I'm running on threads %s" % threading.current_thread())
print(f"I'm running on threads {threading.current_thread()}")

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function job1 refactored with the following changes:

Comment threadSchedule/demo8.py
print(f"I'm running on threads {threading.current_thread()}")
def job2():
print("I'm running on threads %s" % threading.current_thread())
print(f"I'm running on threads {threading.current_thread()}")

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function job2 refactored with the following changes:

Comment threadSchedule/demo8.py
print(f"I'm running on threads {threading.current_thread()}")
def job3():
print("I'm running on threads %s" % threading.current_thread())
print(f"I'm running on threads {threading.current_thread()}")

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function job3 refactored with the following changes:

Comment on lines -13 to +15
CHROME_DRIVER_MAPPING_FILE = r"{}\mapping.json".format(CHROME_DRIVER_FOLDER)
CHROME_DRIVER_EXE = r"{}\chromedriver.exe".format(CHROME_DRIVER_FOLDER)
CHROME_DRIVER_ZIP = r"{}\chromedriver_win32.zip".format(CHROME_DRIVER_FOLDER)
CHROME_DRIVER_MAPPING_FILE = f"{CHROME_DRIVER_FOLDER}\mapping.json"
CHROME_DRIVER_EXE = f"{CHROME_DRIVER_FOLDER}\chromedriver.exe"
CHROME_DRIVER_ZIP = f"{CHROME_DRIVER_FOLDER}\chromedriver_win32.zip"

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Lines 13-15 refactored with the following changes:

Comment on lines -386 to +371
ret = []
for i in range(1, page):
ret.append('http://www.samair.ru/proxy/proxy-%(num)02d.htm' % {'num': i})
return ret
return [
'http://www.samair.ru/proxy/proxy-%(num)02d.htm' % {'num': i}
for i in range(1, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_3 refactored with the following changes:

Comment on lines -406 to +388
ip = match[0] + "." + match[1] + match[2]
ip = f"{match[0]}.{match[1]}{match[2]}"

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function parse_page_3 refactored with the following changes:

Comment on lines -432 to +417
ret = []
for i in range(1, page):
ret.append('http://www.pass-e.com/proxy/index.php?page=%(n)01d' % {'n': i})
return ret
return [
'http://www.pass-e.com/proxy/index.php?page=%(n)01d' % {'n': i}
for i in range(1, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_4 refactored with the following changes:

Comment on lines -476 to +461
ret = []
for i in range(1, page):
ret.append('http://www.ipfree.cn/index2.asp?page=%(num)01d' % {'num': i})
return ret
return [
'http://www.ipfree.cn/index2.asp?page=%(num)01d' % {'num': i}
for i in range(1, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_5 refactored with the following changes:

Comment on lines -508 to +493
ret = []
for i in range(1, page):
ret.append('http://www.cnproxy.com/proxy%(num)01d.html' % {'num': i})
return ret
return [
'http://www.cnproxy.com/proxy%(num)01d.html' % {'num': i}
for i in range(1, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_6 refactored with the following changes:

Comment on lines +507 to -528
type = -1 # 该网站未提供代理服务器类型
for match in matches:
ip = match[0]
port = match[1]
type = -1 # 该网站未提供代理服务器类型

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function parse_page_6 refactored with the following changes:

Comment on lines -595 to +580
ret = []
for i in range(0, page):
ret.append('http://proxylist.sakura.ne.jp/index.htm?pages=%(n)01d' % {'n': i})
return ret
return [
'http://proxylist.sakura.ne.jp/index.htm?pages=%(n)01d' % {'n': i}
for i in range(0, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_9 refactored with the following changes:

Comment on lines -616 to +598
if (type == 'Anonymous'):
type = 1
else:
type = -1
type = 1 if (type == 'Anonymous') else -1

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function parse_page_9 refactored with the following changes:

Comment on lines -634 to +616
ret = []
for i in range(1, page):
ret.append('http://www.publicproxyservers.com/page%(n)01d.html' % {'n': i})
return ret
return [
'http://www.publicproxyservers.com/page%(n)01d.html' % {'n': i}
for i in range(1, page)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_10 refactored with the following changes:

Comment on lines -678 to +667
ret = []
for i in range(1, page):
ret.append('http://www.my-proxy.com/list/proxy.php?list=%(n)01d' % {'n': i})

ret.append('http://www.my-proxy.com/list/proxy.php?list=s1')
ret.append('http://www.my-proxy.com/list/proxy.php?list=s2')
ret.append('http://www.my-proxy.com/list/proxy.php?list=s3')
ret = [
'http://www.my-proxy.com/list/proxy.php?list=%(n)01d' % {'n': i}
for i in range(1, page)
]
ret.extend(
(
'http://www.my-proxy.com/list/proxy.php?list=s1',
'http://www.my-proxy.com/list/proxy.php?list=s2',
'http://www.my-proxy.com/list/proxy.php?list=s3',
)
)

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function build_list_urls_11 refactored with the following changes:

urls = []
for page in range(beg, end):
urls.append('http://www.baidu.com?&page=%d' % page)
urls = ['http://www.baidu.com?&page=%d' % page for page in range(beg, end)]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function run_spider refactored with the following changes:

Comment on lines -92 to +90
res = []
for i in range(len(s) - 1):
res.append((s[i], s[i + 1]))
res = [(s[i], s[i + 1]) for i in range(len(s) - 1)]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function main refactored with the following changes:

Comment on lines -64 to -68
proxies = {
'http': '{proxy_type}://{ip}:{port}'.format(proxy_type=proxy_type, ip=ip, port=port),
'https': '{proxy_type}://{ip}:{port}'.format(proxy_type=proxy_type, ip=ip, port=port),
return {
'http': '{proxy_type}://{ip}:{port}'.format(
proxy_type=proxy_type, ip=ip, port=port
),
'https': '{proxy_type}://{ip}:{port}'.format(
proxy_type=proxy_type, ip=ip, port=port
),
}
return proxies

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function get_proxy_dict refactored with the following changes:

urls = []
for page in range(1, 1000):
urls.append('http://www.jb51.net/article/%s.htm' % page)
urls = [f'http://www.jb51.net/article/{page}.htm' for page in range(1, 1000)]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function main refactored with the following changes:

Comment threadcrawler/src/test.py
class TestSyncSpider(SyncSpider):
def handle_html(self, url, html):
print(html)
pass

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function TestSyncSpider.handle_html refactored with the following changes:

Comment threadcrawler/src/test.py
class TestAsyncSpider(AsyncSpider):
def handle_html(self, url, html):
print(html)
pass

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function TestAsyncSpider.handle_html refactored with the following changes:

Comment threadcrawler/src/test.py
Comment on lines -24 to +25
urls = []
for page in range(1, 1000):
#urls.append('http://www.jb51.net/article/%s.htm' % page)
urls.append('http://www.imooc.com/data/check_f.php?page=%d'%page)
urls = [
'http://www.imooc.com/data/check_f.php?page=%d' % page
for page in range(1, 1000)
]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Lines 24-27 refactored with the following changes:

This removes the following comments ( why? ):

#urls.append('http://www.jb51.net/article/%s.htm' % page)

Comment on lines -64 to +69
func_name = 'get_' + field
func_name = f'get_{field}'
xpath_str = self.xpath_dict.get(field)
if hasattr(self, func_name):
return getattr(self, func_name)(xpath_str)
else:
self.logger.debug(field, self.url)
return self.parser.xpath(xpath_str)[0].strip() if xpath_str else ''
self.logger.debug(field, self.url)
return self.parser.xpath(xpath_str)[0].strip() if xpath_str else ''

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function XpathCrawler.get_field refactored with the following changes:

Comment on lines -73 to +75
xpath_result = {}
for field, xpath_string in self.xpath_dict.items():
xpath_result[field] = makes(self.get_field(field)) # to utf8
return xpath_result
return {
field: makes(self.get_field(field))
for field, xpath_string in self.xpath_dict.items()
}

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function XpathCrawler.get_result refactored with the following changes:

This removes the following comments ( why? ):

# to utf8

Comment on lines -93 to +92
category_urls = []
for href in category_hrefs:
category_urls.append(urljoin(self.domain, href))
category_urls = [urljoin(self.domain, href) for href in category_hrefs]

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function FantasyhairbuySite.generate_category_urls refactored with the following changes:

Comment on lines -11 to +14
backup_name = os.path.basename(file_path) + '_' + datetime.datetime.now().strftime('%Y%m%d%H%M%S')
backup_name = (
f'{os.path.basename(file_path)}_'
+ datetime.datetime.now().strftime('%Y%m%d%H%M%S')
)

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function backup_file refactored with the following changes:

Comment on lines -44 to +47
backup_list = glob.glob(os.path.join(path, pattern + '_*'))
backup_list = glob.glob(os.path.join(path, f'{pattern}_*'))

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function main refactored with the following changes:


# if it's anything else, return it in its original form
return data
return data.encode('utf-8') if str(type(data)) == "<type 'unicode'>" else data

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function _byteify refactored with the following changes:

This removes the following comments ( why? ):

# if it's anything else, return it in its original form

try:
res = query.find()
return res
return query.find()

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function LeanCloudApi.get_skip_obj_list refactored with the following changes:

Comment on lines -49 to +48
img_info_url = img_url + '?imageInfo'
img_info_url = f'{img_url}?imageInfo'

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Function LeanCloudApi.add_img_info refactored with the following changes:

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants