Reading HTTP Server Streams with Python

I'm playing around trying to write a client for a site that provides data as an HTTP stream (also known as an HTTP server). However urllib2.urlopen () grabs the stream in its current state and then closes the connection. I've tried skipping urllib2 and using httplib directly, but it looks like the same behavior.

The request is a POST request with a set of five parameters. However, cookie and authentication are not required.

Is there a way to force the stream to stay open so that it can be checked every program loop for new content rather than waiting for the whole thing to be reloaded every few seconds by typing lag?

+2


a source to share


3 answers


Do you need to actually parse the response titles, or are you primarily interested in content? And your complex HTTP requests for you to set cookies and other headers, or would a simple request be enough?

If you only care about the body of the HTTP response and don't have a very fancy request, you should just use a socket connection:



import socket

SERVER_ADDR = ("example.com", 80)

sock = socket.create_connection(SERVER_ADDR)
f = sock.makefile("r+", bufsize=0)

f.write("GET / HTTP/1.0\r\n"
      + "Host: example.com\r\n"    # you can put other headers here too
      + "\r\n")

# skip headers
while f.readline() != "\r\n":
    pass

# keep reading forever
while True:
    line = f.readline()     # blocks until more data is available
    if not line:
        break               # we ran out of data!

    print line

sock.close()

      

+1


a source


You can try lib queries.

import requests
r = requests.get('http://httpbin.org/stream/20', stream=True)

    for line in r.iter_lines():

    # filter out keep-alive new lines
    if line:
        print line

      



You can also add parameters:

import requests
settings = { 'interval': '1000', 'count':'50' }
url = 'http://agent.mtconnect.org/sample'

r = requests.get(url, params=settings, stream=True)

for line in r.iter_lines():
    if line:
        print line

      

+1


a source


One way to do it with urllib2

(if this site also requires Basic Auth):

 import urllib2
 p_mgr = urllib2.HTTPPasswordMgrWithDefaultRealm()
 url = 'http://streamingsite.com'
 p_mgr.add_password(None, url, 'login', 'password')

 auth = urllib2.HTTPBasicAuthHandler(p_mgr)
 opener = urllib2.build_opener(auth)

 urllib2.install_opener(opener)
 f = opener.open('http://streamingsite.com')

 while True:
     data = f.readline()

      

0


a source







All Articles