Tuesday, January 6, 2009

Writing webbots using Python

If you ever wanted to write a webbot and didn't know how, it could be easily achieved using the mechanize module.

There is one thing that mechanize does which doesn't always suit my needs, which is to pay attention to the robots.txt file, so I just disable it.

Mechanize allows you to easily browse, extract data and submit forms.

import mechanize

br = mechanize.Browser()
br.set_handle_robots(False)
br.open("http://www.google.com")

for link in br.links():
print link

Using Python and MySQL

I use the MySQLPython module, which is very easy to use.
To return results as a dictionary or as a tuple, we use the DictCursor to fetch rows.

import MySQLdb
import MySQLdb.cursors

conn = MySQLdb.connect(
host = "localhost",
user = "xxx",
passwd = "xxx",
db = "xxx",
cursorclass = MySQLdb.cursors.DictCursor)

cur = conn.cursor()

cur.execute("select id, address from phonebook where city_id = 0")

data = cur.fetchall()

Wednesday, December 3, 2008

GeoIP MySQL database creation and CSV loading script

If you need to create a GeoIPCity database and load the CSV files from GeoIP, just run this SQL script (make sure you put the CSV files in the right place).

After loading the DB, just run the following SQL command:

SELECT * FROM geoip_blocks JOIN geoip_loc ON geoip_blocks.locId = geoip_loc.locId
WHERE BETWEEN startIpNum AND endIpNum


CREATE TABLE `geoip_blocks` (
`startIpNum` BIGINT NOT NULL ,
`endIpNum` BIGINT NOT NULL ,
`locId` BIGINT NOT NULL
) ENGINE = MYISAM CHARACTER SET utf8 COLLATE utf8_bin;

CREATE TABLE `geoip_loc` (
`locId` BIGINT NOT NULL ,
`country` VARCHAR( 2 ) NULL ,
`region` VARCHAR( 3 ) NULL ,
`city` VARCHAR( 100 ) NULL ,
`postalCode` VARCHAR( 10 ) NULL ,
`latitude` FLOAT NOT NULL ,
`longitude` FLOAT NOT NULL ,
`metroCode` INT NULL ,
`areaCode` INT NULL ,
PRIMARY KEY ( `locId` )
) ENGINE = MYISAM CHARACTER SET utf8 COLLATE utf8_bin;

LOAD DATA INFILE '/root/geoip/GeoLiteCity-Blocks.csv' INTO TABLE geoip_blocks FIELDS OPTIONALLY ENCLOSED BY '"' TERMINATED BY ',' IGNORE 2 LINES;

LOAD DATA INFILE '/root/geoip/GeoLiteCity-Location.csv' INTO TABLE geoip_loc FIELDS OPTIONALLY ENCLOSED BY '"' TERMINATED BY ',' IGNORE 2 LINES;

Tuesday, April 24, 2007

DDoS using XSS and Ajax

A thought I had a few days ago about Distributed Denial of Service...

DDoS is usually obtained using a botnet that receives a command to enter a certain website at once to choke its bandwidth. This kind of attack is almost unstoppable since there is usually no way of knowing who are the legitimate users and what page requests came from bots.
But what if someone found a way to run a JS script using XSS on a very big website with tens of thousands of hits per day? What if that script contained a small deferred background JS script that continuously creates simple XMLHTTP requests to a certain page?

Monday, April 23, 2007

Cheating on Bandwidth with PHP

Note: I am NOT responsible for anything that could happen if you actually try this (especially if your web hosting service decides to sue you).

PHP is a very fun scripting language. Besides the stanard behaviour that is expected from an honest script to connect to a database, parse the information and display it to the user, PHP can do a few tricks as well. One of them, is to create a listening socket and forking a new process.

If your web hosting allows you to fork and create a listening socket through PHP, you might just be able to do some nasty things so that your visitors will download content from another port on the server, instead of through the Apache web server. Traffic that is being downloaded from your site through Apache gets summed up and limited, usually for a fixed amount per month, depending on your hosting package. If you decide to get more bandwidth, you need to pay more.

But what if you could get your site's visitors to download the content itself from a forked process from PHP that listens to a specific port and acts like a web server? The user will manage to downlaod content from your site, and the bandwidth will not be accounted for.

The algorithm:
  1. Use mod_rewrite on specific directories to send the requested file name to send through a PHP script
  2. The PHP script will either create the download link or send a Redirect header in the following manner: http://www.yoursite.com:[random_port], while random_port is a number between 1024 and 65535 (Ports below 1024 are privileged ports).
  3. At the same time that the link is created, use fork to create a small temporary daemon that will run in the background and wait for a connection.
  4. The client will attempt to download the file through the chosen port.
  5. The forked PHP script will parse the HTTP request and send the requested file back to the client. The HTTP request will be a very simple one (something like "GET / HTTP/1.1" and a few more insignificant headers) since we already know exactly which file to send. (One of the parameters to our script was the file name).
  6. It is possible to leave the forked daemon on, but that would really be nasty :)

Probably the most obvieous reason why this won't work usually is because of PHP security settings that will not allow you to do this hack. Other than that, most servers today have firewalls for incoming connections, especially on unprivileged ports. If your server's hosting is lame enough, you might actually succeed in doing this.

WYSIWYGS Edtiros Suck, Use WYSIWYM Editors

The basic problem today with developing or using Content Management Systems, Blogs, etc. is that we want to allow the content writers to have a flexible editor with features such as bullets, bold font, different font size, etc. On the other hand, we want to have a strict CSS design for our website, so that the stylesheet will determine how the content will be displayed in a unified manner. With WYSIWYM editors, both goals can be acheived, since WYSIWYM editors generate strict and standard XHTML code, which was designed specifically for this purpose.

Anyways, here is an excellent article about why WYSIWYM editors kick ass:

http://www.456bereastreet.com/archive/200612/forget_wysiwyg_editors_use_wysiwym_instead/

Friday, April 20, 2007

Web 2.0 - Beware!

The new AJAX approach to web design is fun and fascinating, but dangerous at the same time. The main problem with AJAX is that you can't index your site easily. If most of your website content is generated dynamically in the page using AJAX, search engines will NOT be able to index your site content. This issue is supposed to be figured out sometimes, and I'm sure Google is already working on a Javascript / browser simulator to solve this issue out. But Google's solution to dynamic content will never be perfect, because Web 2.0 usually relies on human interaction.

The more concerning issue about Web 2.0 is content stealing - since the basic idea behind AJAX requests is client side data processing (which gives the web much more flexability), the data that is received at the client is plaintext and can usually be parsed in a simple manner (XML or CSV data). The problem is that it becomes very easy to reverse engineer AJAX driven webpages because of the low security implementation. It is much harder to reverse engineer a program and understand how it connects to its remote server, or parse data by yourself from server side web applications. Stealing a webpage written with AJAX can be as simple as copy-pasting functions from the original web page.

So how can these application be protected?

First of all, obfuscation of the data and the code itself. There are program that know how to do it and it might be very helpful to defend against the most common and lamest hackers around. Data obfuscation can be obtained by a simple encryption which is hard to understand and easy to process using Javascript.

The data source itself can be also protected using a referrer check - if the AJAX request came from an unknown page, the service can be blocked. But this can also be easily bypassed by forging the referrer header from the client or from servers that rip the data from the service.

The best technique for protecting AJAX services is using a session - either by using login cookies which the AJAX requests use, or server generated random values that pass back manually from the Javascript itself (the exact same idea, only does not need cookie support and a bit harder to implement). This method is the exact same technique that is used to protect sites from unauthorized users, only that the login sequence is automatic once you enter the main page.

Of course that temporary session cookies are not enough to protect AJAX sites, since another request can be added to extract the session cookie from the main page automatically from the client, which is usually a difficult task to do, exactly as difficult as ripping sites would be, which is exactly what we wanted to achieve.