Wednesday, August 31, 2011

The trouble with Protected Mode in Selenium 2 for setting cookies...

If you try to set cookies using the WebDriver API through Selenium 2, you may find that Internet Explorer fails to even set the cookie. The issue has been reported here:

http://code.google.com/p/selenium/issues/detail?id=1227&q=cookie&colspec=ID%20Stars%20Type%20Status%20Priority%20Milestone%20Owner%20Summary#makechanges

Selenium has a bunch of test suites to verify the behavior so it seemed strange that there would be an issue. Examining the IEDriver code too shows that the AddCookieCommandHandler is very similar to behavior of DeleteAllCookiesHandler and other IE command handlers, so I didn't really find an issue. Nor did the AddCookie() method in DocumentHost.cpp handles the dispatching of adding cookies.

If we do:
>>> driver = webdriver.Remote(desired_capabilities=current_env,
                                        command_executor="http://localhost:4444/wd/hb")
>>> driver.execute_script("document.cookie='a=1';")
>>> driver.get_cookies()
[{u'name': u'a', u'value': u'1', u'path': u'/', u'hCode': 97, u'class': u'org.openqa.selenium.Cookie', u'secure': False}, {u'name': u'sessionid', u'value': u'e1257265399f35b5c7ae4cf630581c90', u'path': u'/', u'hCode': 607797809, u'class': u'org.openqa.selenium.Cookie', u'secure': False}]
>>> driver.delete_cookie('a')
>>> driver.get_cookies()
[{u'name': u'sessionid', u'value': u'e1257265399f35b5c7ae4cf630581c90', u'path': u'/', u'hCode': 607797809, u'class': u'org.openqa.selenium.Cookie', u'secure': False}]
>>> driver.add_cookie({'name' : 'a' , 'value' : '2', 'secure' : False})
(Pdb) driver.get_cookies()
[{u'name': u'sessionid', u'value': u'e1257265399f35b5c7ae4cf630581c90', u'path': u'/', u'hCode': 607797809, u'class': u'org.openqa.selenium.Cookie', u'secure': False}]

Both delete commands will work successfully. But doing an add_cookie() fails to work if IE7/IE8 are in Protected mode in Selenium 2.5.0, even though Internet Explorer/IEDriver does not report any error. However, I was able to get cookies to be set once I disabled Protected mode.

Since Selenium 2.5.0 requires all Protected Mode settings to be consistent, you have to go into your Internet Options and uncheck the Protected Mode for every single zone (i.e. click through the icons for Internet, Local Internet, Trusted sites, Restricted sites). Then cookies can be correctly set.

It appears that IE7 and IE8 have this issue. IE9 may not have this problem. I was not able to get cookies set in either Protected/non-Protected mode using Selenium v2.0.0b3 though so you may still need to upgrade to Selenium 2 beyond this version to get cookie support working in IE7/IE8.

Tuesday, August 30, 2011

Using add_cookie in Selenium 2

The documentation in the Python bindings for using the add_cookie() function Selenium 2 are unclear. The add_cookie() appears to take in a simple key/value pair:

def add_cookie(self, cookie_dict):
        """Adds a cookie to your current session.
        Args:
            cookie_dict: A dictionary object, with the desired cookie name as the key, and
            the value being the desired contents.
        Usage:
            driver.add_cookie({'foo': 'bar',})
        """
        self.execute(Command.ADD_COOKIE, {'cookie': cookie_dict})
If you're encountering NullPointerExceptions similar to a bug it's possible the problem is that your dictionary needs to include name, value, path, and secure keys. The tests in selenium/webdriver/common/cookies_test.py appear to back this point up:

self.COOKIE_A = {"name": "foo",
                         "value": "bar",
                         "path": "/",
                         "secure": False}

    def testAddCookie(self):
        self.driver.execute_script("return document.cookie")
        self.driver.add_cookie(self.COOKIE_A)

Even the section posted at http://readthedocs.org/docs/selenium-python/en/latest/navigating.html#cookies suggest that adding cookie just a matter of connecting to using a key/value pair too:

Before we leave these next steps, you may be interested in understanding how to use cookies. First of all, you need to be on the domain that the cookie will be valid for:

# Go to the correct domain
driver.get("http://www.example.com")

# Now set the cookie. This one's valid for the entire domain
cookie = {"key": "value"})
driver.add_cookie(cookie)

# And now output all the available cookies for the current URL
all_cookies = driver.get_cookies()
for cookie_name, cookie_value in all_cookies.items():
    print "%s -> %s", cookie_name, cookie_value

For disabling the Django debug toolbar in Selenium 2, then the command should be:
self.selenium.add_cookie({"name" : "djdt",
                          "value" : "true",
                          "path" : "/",
                          "secure" : False})

As of Selenium v2.5.0, It appears that all name/value and secure must be specified to avoid triggering the NullPointerException error.

An issue report has been filed here:

http://code.google.com/p/selenium/issues/detail?id=2367

Sunday, August 28, 2011

Extracting audio clips from YouTube videos

If you use the stock Ubuntu v10.04 youtube-dl version, you may encounter this error message when trying to download a YouTube clip:
youtube-dl http://www.youtube.com/watch?v=
ERROR: no fmt_url_map or conn information found in video info
The solution is to git clone the youtube-dl repo and use the latest youtube-dl version:
git clone https://github.com/rg3/youtube-dl.git

Extracting only the audio portion means that you should set the -vn option, which disables video encoding. The -acodec option determines the output format. So you would execute the command (depending if you want Ogg bitstream or Mp3 format)

ffmpeg -i  -vn -acodec vorbis 
ffmpeg -i  -vn -acodec mp3 

Saturday, August 27, 2011

Getting branch support to work with Nose/Coverage

While Ned Batchelder's coverage utility has supported branch measurements for sometime, it hasn't been supported in the main line of the nose unit discovery util. We can see in the upcoming v1.1.3 release that the --cover-branches option will be supported:

http://readthedocs.org/docs/nose/en/latest/news.html

Currently nose v1.1.3 is still labeled as a development version, so you'd have to get pip install nose==dev in order to install this copy.

pip install --upgrade coverage
pip install --upgrade nosexcover
pip install --upgrade nose==dev (1.1.3)

An alternative, which has long been suggested in discussion groups, is to create a .coveragerc file to enable branch coverage by default. This file must be placed in the location where coverage.py is run, not necessarily in your home directory:
[run] 
     branch=True 
If you've enabled things correctly, you should see the header (instead of the default) as follows:
Name                                              Stmts   Miss Branch BrPart  Cover   Missing
---------------------------------------------------------------------------------------------
If you're also using --with-xunit to generate Cobertura-style XML reports, hopefully you should also see the branch conditionals also being tallied correctly too!

Wednesday, August 17, 2011

Facebook's OAuth2 support for Python

Facebook recently announced that they will be phasing in OAuth 2.0 support and require its use starting October 1, 2011. On the JavaScript SDK side, there are several  changes on the JavaScript code that have to be done, which are listed as follows.

You can download the Python code here:

https://github.com/rogerhu/facebook_oauth2

1. FB.init has to be initialized with the Facebook APP ID instead of the API Key, though the apiKey parameter still is used.

2. The oauth: true options. must be set in the FB.init() calls.

In other words, the code changes would be:
FB.init({apiKey: facebook_app_id,
             oauth: true,
             cookie: true});
3. Instead of response.session, the response should now be response.authResponse. Also,
make note that scope: should be used instead of perms:
FB.login(function(response) {
    if (response.authResponse) {
    },
    {scope: 'email,publish_stream,manage_pages'}
    });
Also, if you need to retrieve the user id on the JavaScript, the value is stored as response.authResponse.userID instead of response.session.uid:
FB.api(
       { method: 'fql.query',
        query: 'SELECT ' + permissions.join() + ' FROM permissions WHERE uid=' + response.authResponse.userID},
        function (response) { });

If you see yourself not being able to logout, it means you haven't set the right APP ID or forgot to set oauth: true in both your login and logout code. If you're going to make the change, you should make it everywhere in your code!

On the Python/Django side, you need to implement a few helper routines. If Facebook authenticates properly, a cookie with the prefix fbsr_ will be set as a cookie (instead of fbs_). This signed request includes an encoded signature and payload, which must be separated and verified. You can look at the PHP SDK code to understand how it's implemented, or you can review this Python version of the code (see http://developers.facebook.com/docs/authentication/signed_request/)
def parse_signed_request(signed_request, secret):

    encoded_sig, payload = signed_request.split('.', 2)

    sig = base64_urldecode(encoded_sig)
    data = json.loads(base64_urldecode(payload))

    if data.get('algorithm').upper() != 'HMAC-SHA256':
        return None
    else:
        expected_sig = hmac.new(secret, msg=payload, digestmod=hashlib.sha256).digest()

    if sig != expected_sig:
        return None

    return data
In the PHP SDK code, there is a base64_url_decode function that automatically adds the correct number of "=" characters to the end of the Base64 encoded string. The basic problem is that Base64 encodes 3 bytes for every 4 characters, so the total length will be 4*len(string)/3. We can use this knowledge to realize that the total length will be a multiple of 4 and then insert the appropriate number of '=' characters to the end of the string. Facebook also appears to use a Base64-uRL variant in which the '+' and '/' characters of standard Base64 are respectively replaced by '-' and '_', which then must be replaced during the decode process (see http://en.wikipedia.org/wiki/Base64#URL_applications). The code looks like the following:
def base64_urldecode(data):
    # http://qugstart.com/blog/ruby-and-rails/facebook-base64-url-decode-for-signed_request/                       
    # 1. Pad the encoded string with "+".                                                                          
    # See http://fi.am/entry/urlsafe-base64-encodingdecoding-in-two-lines/                                         
    data += "=" * (4 - (len(data) % 4) % 4)

    return base64.urlsafe_b64decode(data)
If you're using the old Python SDK implementation, you may wish to implement code that mimics the way in which the Python SDK implemented get_user_from_cookie, since the expires, session_key, and oauth_token can be derived from retrieving the access token. We also set an fbsr_signed parameter in case you have debugging statements in your code and want to differentiate between your old get_user_from_cookie from this code.

Note: in order to make things backward-compatible, you need to make an extra URL request back to Facebook to retrieve the access token. This code was also inspired from the Facebook PHP SDK code too:
def get_access_token_from_code(code, redirect_url=None):
    """ OAuth2 code to retrieve an application access token. """

    data = {
        'client_id' : settings.FACEBOOK_APP_ID,
        'client_secret' : settings.FACEBOOK_SECRET_KEY,
        'code' : code,
        }

    if redirect_url:
        data['redirect_uri'] = redirect_url
    else:
        data['redirect_uri'] = ''


   return get_app_token_helper(data)

BASE_LINK = "https://graph.facebook.com"

def get_app_token_helper(data=None):
   
    if not data:
        data = {}

    try:
        token_request = urllib.urlencode(data)

        app_token = urllib2.urlopen(BASE_LINK + "/oauth/access_token?%s" % token_request).read()
    except urllib2.HTTPError, e:
        logging.debug("Exception trying to grab Facebook App token (%s)" % e)
        return None

    matches = re.match(r"access_token=(?P.*)", app_token).groupdict()

    return matches.get('token')

Tuesday, August 16, 2011

How to redirect bash time outputs..

http://mywiki.wooledge.org/BashFAQ/032


How can I redirect the output of 'time' to a variable or file?

Bash's time keyword uses special trickery, so that you can do things like
   time find ... | xargs ...
and get the execution time of the entire pipeline, rather than just the simple command at the start of the pipe. (This is different from the behavior of the external command time(1), for obvious reasons.)
Because of this, people who want to redirect time's output often encounter difficulty figuring out where all the file descriptors are going. It's not as hard as most people think, though -- the trick is to call time in a SubShell or block, and then capture stderr of the subshell or block (which will contain time's results). If you need to redirect the actual command's stdout or stderr, you do that inside the subshell/block. For example:
  • File redirection:
       bash -c "time ls" 2>time.output      # Explicit, but inefficient.
       ( time ls ) 2>time.output            # Slightly more efficient.
       { time ls; } 2>time.output           # Most efficient.
    
       # The general case:
       { time some command >stdout 2>stderr; } 2>time.output