Friday, June 17, 2011

Making a message box in Android

Turns out that making a simple message box in Android is just complicated. So here's a very useful snippet:


public void messageBox(String title, String message) {
AlertDialog alertDialog;
alertDialog = new AlertDialog.Builder(MyActivity.this).create();
alertDialog.setTitle(title);
alertDialog.setMessage(message);
alertDialog.setButton("OK", new DialogInterface.OnClickListener() {
@Override
public void onClick(DialogInterface dialog, int which) {

dialog.cancel();

}
});
alertDialog.show();

}

Friday, May 27, 2011

Why Microsoft is going to dominate 50% of the mobile world in about 5 years

Microsoft nowadays looks like it's desperately trying to hold a grip on something. The recent acquisition of Skype, Windows Phone 7, and other things just come to show us how lost they are. They are even planning ARM support for Windows 8, something which they tried to avoid for a long time. Since the uprising power of the ARM processors, alongside the amazing power consumption, it just looks like their way out of the mud. It seems that they are decades away from Apple and Google, but they do not even realize how far ahead they are.

The Windows x86 platform is the most common operating system in the world today. Therefore, trillions of code lines in the world have been written and adjusted for this platform, and Windows 7 continues to be the primary, and probably most successful operating system today, especially in the enterprise world.

As of March 2011, more than 80% of the world's consumer devices are Windows powered - XP, Vista and 7. But how can this advantage be leveraged into the mobile world against Android and iOS? You'd obviously think that an ARM compatible, slim embedded operating system from Microsoft is the solution, and so do Microsoft. That's why Windows 8 is planned to be compatible with ARM. But the answer is much closer than anyone would expect.

The battle between ARM and Intel is emerging - ARM, from one side is trying to breach the high end CPU market, whereas Intel, using their Atom CPU are trying to compete in the mobile market. But up to now, Intel's new System-on-Chip - code named Silvermont, using their new Tri-gate transistor technology cuts down the power consumption in half, which makes it just good enough to manufacture smart-phones. By 2013, very power efficient smart-phones will be emerging with enough processing power and memory to run a fully fledged Windows 7 operating system, which is exactly what they need to leverage their advantage in the PC market. Once Windows 7 mobile phones will be available, they will start replacing desktops and laptops using docking stations and mobile modules, utilizing a mobile interface when on the go, or a normal desktop when docked. This use case is perfect for enterprise, where people need their laptops with them at all times, wherever they are, and would love saving the hassle of carrying around their laptop around.

The many users who use Android and iOS today, who still use a Windows based operating system, would prefer a Windows mobile phone, if they were good enough. But so far, Microsoft is losing this battle, and losing hard. Windows Phone 7, just like Android 2 and iOS will become obsolete, as Android 3 or Chrome OS from Google, OS X for mac and Windows 7/8 from Microsoft will take over the mobile market as well, once power efficient, small factor Atom CPUs will emerge, like Silvermont. The mobile operating systems will become history. Once this happens, the current operating system market share will start reflecting on the mobile devices, and will increase Microsoft's share to become at least 50% (and even up to 80%) of the mobile market.

The upcoming Windows 8 tablets and other mobile devices, based on ARM, which will probably start to emerge in the upcoming year is a step forward in terms of Windows 8, but the ARM support could only help Microsoft in taking over a larger share of the market. However, this is not the reason for their upcoming success - their success will emerge from the fact that Windows is the most commonly used OS today, running on the x86 platform.

Microsoft still has a long way to go. But the fact is, that their natural advantage will allow them to succeed in the mobile market, without the need to work hard, like Apple and Google. By 2016 or so, mobile platforms will be as fast as modern PCs, allowing everyone to have their handheld computer in their hands at all times, playing CPU intense computer games they recently installed, and using Outlook, Office and Visual studio at the same time.




Friday, July 16, 2010

How to root an HTC Desire using Unrevoked on a 64 bit Windows

Just recently, the "Unrevoked" one click rooting program became available from http://unrevoked.com.

The "Unrevoked" root won't run on a Windows 7 64 bit out of the box, on an HTC desire, because of a missing driver. This can be easily fixed in a few steps.

Some of you might have noticed that neither the HTC sync drivers nor the Android SDK have the "Android Bootloader Interface" drivers for a 64 bit Windows. You can't really get them either.

Luckily, the Android SDK drivers have a 64 bit version of the ABI driver. The problem is, that the USB identification of the device is unknown to the driver. To fix it, we need to add the following line to the "android_winusb.inf" file:

[Google.NTamd64]
%SingleBootLoaderInterface% = USB_Install, USB\VID_0BB4&PID_0C94&REV_0100

This line defines that on a 64 bit machine, if the USB device with a vendor ID of 0x0BB4 (HTC) and a product ID of 0x0C94 (Bootloader interface on HTC Desire), then the driver that should be installed is "%SingleBootLoaderInterface%".

After that, install the driver you altered and that's it.

This is how the USB driver identifies itself (when the driver is not installed, you can look at the hardware IDs of the device identified as "Android 1.0" (which is the bootloader interface before the driver is installed).
Just right click the driver, click Properties, go to Details, and look at the "Hardware Ids" section. You can do the same trick for any other Android device which supports ABI.

Friday, April 23, 2010

Convenient thread logging in Python

I wrote a small logging thread class for easy logging.
My standard error now looks like this:
[Fri Apr 23 15:58:20 2010] [ClipboardReader-1] Starting log...
[Fri Apr 23 15:58:20 2010] [SocksValidator-1] Starting log...
[Fri Apr 23 15:58:20 2010] [SocksChecker-1] Starting log...
Every thread also has its own logfile in the "logs" directory by his class name. I made up the thread name to add the class name of the instance (won't be LoggingThread if you inherit from it).
class LoggingThread(threading.Thread):
def __init__(self, *args, **kwargs):
threading.Thread.__init__(self, *args, **kwargs)
if hasattr(self.__class__, "instance_count"):
self.__class__.instance_count += 1
else:
self.__class__.instance_count = 1

self.name = "%s-%d" % (self.__class__.__name__, self.__class__.instance_count)

if not os.path.isdir("logs"):
os.makedirs("logs")
self.logfile = open("logs/%s.log" % self.name, "w")
self.log("Starting log...")

def log(self, data):
logline = "[%s] [%s] %s" % (time.ctime(), self.name, data)
print >> self.logfile, logline
self.logfile.flush()
print >> sys.stdout, logline

Getting and setting text from the clipboard using Python

Just name this module clipboard.py and use GetClipboardText/SetClipboardText. Very handy, and doesn't really need the win32 extentions, although it's used here.


from ctypes import *
from win32con import CF_TEXT, GHND

OpenClipboard = windll.user32.OpenClipboard
EmptyClipboard = windll.user32.EmptyClipboard
GetClipboardData = windll.user32.GetClipboardData
SetClipboardData = windll.user32.SetClipboardData
CloseClipboard = windll.user32.CloseClipboard
GlobalLock = windll.kernel32.GlobalLock
GlobalAlloc = windll.kernel32.GlobalAlloc
GlobalUnlock = windll.kernel32.GlobalUnlock
memcpy = cdll.msvcrt.memcpy

def GetClipboardText():
text = ""
if OpenClipboard(c_int(0)):
hClipMem = GetClipboardData(c_int(CF_TEXT))
GlobalLock.restype = c_char_p
text = GlobalLock(c_int(hClipMem))
GlobalUnlock(c_int(hClipMem))
CloseClipboard()
return text

def SetClipboardText(text):
buffer = c_buffer(text)
bufferSize = sizeof(buffer)
hGlobalMem = GlobalAlloc(c_int(GHND), c_int(bufferSize))
GlobalLock.restype = c_void_p
lpGlobalMem = GlobalLock(c_int(hGlobalMem))
memcpy(lpGlobalMem, addressof(buffer), c_int(bufferSize))
GlobalUnlock(c_int(hGlobalMem))
if OpenClipboard(0):
EmptyClipboard()
SetClipboardData(c_int(CF_TEXT), c_int(hGlobalMem))
CloseClipboard()

Thursday, April 22, 2010

Here is a really nice piece of code which implements a TCP relay (tunnel) using Python's asyncore module. It's really interesting because the way I implemented it makes a lot of sense, in contrary to the same code written using select.

There's also something about using asyncore.dispatcher_with_send on Unix systems which might send an EWOULDBLOCK sometimes, but I didn't dig deep enough.

The 3 sockets playing a role here are:
* A Relay server socket - accepts connections from clients
* Relay clients, which trigger a new connection to the tunnel destination
* Relay connection for each relay client connected.

The cool thing about asyncore is that you can also decide if handle_read will be called or not, even if data is available, by overriding the "readable" function. In this implementation, a relay client does not start reading from a socket until the relay connection has been established successfully.




class RelayConnection(asyncore.dispatcher):
def __init__(self, client, address):
asyncore.dispatcher.__init__(self)
self.client = client
self.create_socket(socket.AF_INET, socket.SOCK_STREAM)
print "connecting to %s..." % str(address)
self.connect(address)

def handle_connect(self):
print "connected."
# Allow reading once the connection
# on the other side is open.
self.client.is_readable = True

def handle_read(self):
self.client.send(self.recv(1024))

class RelayClient(asyncore.dispatcher):
def __init__(self, server, client, address):
asyncore.dispatcher.__init__(self, client)
self.is_readable = False
self.server = server
self.relay = RelayConnection(self, address)

def handle_read(self):
self.relay.send(self.recv(1024))

def handle_close(self):
print "Closing relay..."
# If the client disconnects, close the
# relay connection as well.
self.relay.close()
self.close()

def readable(self):
return self.is_readable

class RelayServer(asyncore.dispatcher):
def __init__(self, bind_address, dest_address):
asyncore.dispatcher.__init__(self)
self.create_socket(socket.AF_INET, socket.SOCK_STREAM)
self.bind(bind_address)
self.dest_address = dest_address
self.listen(10)

def handle_accept(self):
conn, addr = self.accept()
RelayClient(self, conn, self.dest_address)


RelayServer(("0.0.0.0", 8080), ("127.0.0.1", 1234))
asyncore.loop()

Wednesday, July 29, 2009

Iterating over a sequence of IPs in Python

I was bit bored, so I wrote a simple iterator that iterates over a sequence of IPs using a subnet bitmask.

import struct
import socket

class IPIterator:
def __init__(self, ip, bitmask):
self.ip = ip
self.bitmask = bitmask
def __iter__(self):
ipnum = struct.unpack(">L", socket.inet_aton(self.ip))[0]
mask = 2 ** self.bitmask - 1
for x in xrange(ipnum & ~mask, ipnum | mask):
yield socket.inet_ntoa(struct.pack(">L", x))

if __name__ == "__main__":
for x in IPIterator("1.2.3.4", 8):
print x

Friday, February 6, 2009

Playing with signals in Python

A Python program is able to catch signals and handle them with the signal module. This is especially useful when you want to clean up after your program exits for any reason. This can be done with signal.signal, which sets a callback function for signal handlers.

This code will set an "alarm" that must be set every 30 seconds, or else a signal will be called to shut down the program (sort of a "watchdog" program).


def alarm_handler(signum, frame):
print "30 seconds have passed!"

signal.signal(signal.SIGALRM, alarm_handler)

signal.alarm(30)

while True:
data = raw_input("Write something in the next 30 seconds please: ")
signal.alarm(30)


Another use of signals would be to brutally terminate your program, since sometimes Python can hang on I/O, or a Thread object that is running (for some reason a Ctrl+C doesn't kill a program running a python thread):


import os, signal

# Brutally terminate this process
os.kill(0, signal.SIGTERM)

Sending E-Mail using Python

Well, it's really easy, all you need to do is build up a MIME message using the Python "email" library.

I am using the "alternative" multipart content type, which allows be to enter both text content and HTML content, and display the supported content type (which will of course almost always be HTML, but there are a lot of guys who still use text clients to read emails apparently).

I also used a small line in the example that convert from simple text to HTML.

Here is a sample code:


#!/usr/bin/python
import sys
import smtplib
from email.mime.text import MIMEText
from email.mime.image import MIMEImage
from email.mime.multipart import MIMEMultipart

# We are using the local SMTP server
MAIL_SERVER = "localhost"

class Mailer:
def __init__(self):
self.mail_server = MAIL_SERVER

def send_raw_email(self, mail_from, mail_to, data):
smtp = smtplib.SMTP(self.mail_server)
smtp.sendmail(mail_from, mail_to, data)
smtp.quit()

def send_email(self, mail_from, mail_to, subject, text_message, html_message='', images=[]):
# Initialize MIME message
msg = MIMEMultipart("alternative")
msg["Subject"] = subject
msg["From"] = mail_from
if type(mail_to) is str:
msg["To"] = mail_to
else:
msg["To"] = ", ".join(mail_to)

text = MIMEText(text_message)
msg.attach(text)

if html_message:
html = MIMEText(html_message, "html")
msg.attach(html)

for image in images:
image = MIMEImage(open(image, "rb"))
msg.attach(image)

self.send_raw_email(mail_from, mail_to, msg.as_string())

if __name__ == '__main__':
mailer = Mailer()

subject = "This is the subject"
message = "This is the message"

html = "<html dir="'ltr'"><body>%s</body></html>" % message.replace("\n", "<br/>")
mailer.send_email("admin@myserver.com", "recipient@example.com", subject, message, html)

Thursday, February 5, 2009

Really cheap webhosting?

If you ever wanted to have unlimited websites managed all from one place, instead of paying for extremely limited access to hosting servers which are insecure, and expensive?

I got myself my first VPS a few days ago from http://www.vpsvillage.com. It costs me only 6$/mo and I get to put there as many websites I want (well, it's my server, that's the point of it). Not only that, but a VPS is much more secure than shared webhosting. If a hacker takes over a shared hosting server from any webpage, he could easily take over your files and read your source code from the websites (which could often contain passwords and other secret information). It's not that it's impossible to take over a VPS hosting server, it's just much more difficult to do even if a hacker succeeds on taking over one of the VPS servers.

Of course there's a catch - you have to learn how to be your own webhosting IT manager. I admit - it's not an easy task, but once you learn how to do it, having full control over your website is a huge benefit.

I would suggest using Debian for a start, and reading some HOWTO's on howtoforge.com. For example, to set up a LAMP server (Linux/Apache/MySQL/PHP), what you should do is follow simple instructions from here - http://www.howtoforge.com/ubuntu_debian_lamp_server. Setting up unlimited (duuuh) mailboxes and mail forwarding is easily done using postfix, and you can even set up your own DNS server on the VPS, so adding domains to your VPS is also very easy.

Tuesday, January 6, 2009

Writing webbots using Python

If you ever wanted to write a webbot and didn't know how, it could be easily achieved using the mechanize module.

There is one thing that mechanize does which doesn't always suit my needs, which is to pay attention to the robots.txt file, so I just disable it.

Mechanize allows you to easily browse, extract data and submit forms.

import mechanize

br = mechanize.Browser()
br.set_handle_robots(False)
br.open("http://www.google.com")

for link in br.links():
print link

Using Python and MySQL

I use the MySQLPython module, which is very easy to use.
To return results as a dictionary or as a tuple, we use the DictCursor to fetch rows.

import MySQLdb
import MySQLdb.cursors

conn = MySQLdb.connect(
host = "localhost",
user = "xxx",
passwd = "xxx",
db = "xxx",
cursorclass = MySQLdb.cursors.DictCursor)

cur = conn.cursor()

cur.execute("select id, address from phonebook where city_id = 0")

data = cur.fetchall()

Wednesday, December 3, 2008

GeoIP MySQL database creation and CSV loading script

If you need to create a GeoIPCity database and load the CSV files from GeoIP, just run this SQL script (make sure you put the CSV files in the right place).

After loading the DB, just run the following SQL command:

SELECT * FROM geoip_blocks JOIN geoip_loc ON geoip_blocks.locId = geoip_loc.locId
WHERE BETWEEN startIpNum AND endIpNum


CREATE TABLE `geoip_blocks` (
`startIpNum` BIGINT NOT NULL ,
`endIpNum` BIGINT NOT NULL ,
`locId` BIGINT NOT NULL
) ENGINE = MYISAM CHARACTER SET utf8 COLLATE utf8_bin;

CREATE TABLE `geoip_loc` (
`locId` BIGINT NOT NULL ,
`country` VARCHAR( 2 ) NULL ,
`region` VARCHAR( 3 ) NULL ,
`city` VARCHAR( 100 ) NULL ,
`postalCode` VARCHAR( 10 ) NULL ,
`latitude` FLOAT NOT NULL ,
`longitude` FLOAT NOT NULL ,
`metroCode` INT NULL ,
`areaCode` INT NULL ,
PRIMARY KEY ( `locId` )
) ENGINE = MYISAM CHARACTER SET utf8 COLLATE utf8_bin;

LOAD DATA INFILE '/root/geoip/GeoLiteCity-Blocks.csv' INTO TABLE geoip_blocks FIELDS OPTIONALLY ENCLOSED BY '"' TERMINATED BY ',' IGNORE 2 LINES;

LOAD DATA INFILE '/root/geoip/GeoLiteCity-Location.csv' INTO TABLE geoip_loc FIELDS OPTIONALLY ENCLOSED BY '"' TERMINATED BY ',' IGNORE 2 LINES;

Tuesday, April 24, 2007

DDoS using XSS and Ajax

A thought I had a few days ago about Distributed Denial of Service...

DDoS is usually obtained using a botnet that receives a command to enter a certain website at once to choke its bandwidth. This kind of attack is almost unstoppable since there is usually no way of knowing who are the legitimate users and what page requests came from bots.
But what if someone found a way to run a JS script using XSS on a very big website with tens of thousands of hits per day? What if that script contained a small deferred background JS script that continuously creates simple XMLHTTP requests to a certain page?

Monday, April 23, 2007

Cheating on Bandwidth with PHP

Note: I am NOT responsible for anything that could happen if you actually try this (especially if your web hosting service decides to sue you).

PHP is a very fun scripting language. Besides the stanard behaviour that is expected from an honest script to connect to a database, parse the information and display it to the user, PHP can do a few tricks as well. One of them, is to create a listening socket and forking a new process.

If your web hosting allows you to fork and create a listening socket through PHP, you might just be able to do some nasty things so that your visitors will download content from another port on the server, instead of through the Apache web server. Traffic that is being downloaded from your site through Apache gets summed up and limited, usually for a fixed amount per month, depending on your hosting package. If you decide to get more bandwidth, you need to pay more.

But what if you could get your site's visitors to download the content itself from a forked process from PHP that listens to a specific port and acts like a web server? The user will manage to downlaod content from your site, and the bandwidth will not be accounted for.

The algorithm:
  1. Use mod_rewrite on specific directories to send the requested file name to send through a PHP script
  2. The PHP script will either create the download link or send a Redirect header in the following manner: http://www.yoursite.com:[random_port], while random_port is a number between 1024 and 65535 (Ports below 1024 are privileged ports).
  3. At the same time that the link is created, use fork to create a small temporary daemon that will run in the background and wait for a connection.
  4. The client will attempt to download the file through the chosen port.
  5. The forked PHP script will parse the HTTP request and send the requested file back to the client. The HTTP request will be a very simple one (something like "GET / HTTP/1.1" and a few more insignificant headers) since we already know exactly which file to send. (One of the parameters to our script was the file name).
  6. It is possible to leave the forked daemon on, but that would really be nasty :)

Probably the most obvieous reason why this won't work usually is because of PHP security settings that will not allow you to do this hack. Other than that, most servers today have firewalls for incoming connections, especially on unprivileged ports. If your server's hosting is lame enough, you might actually succeed in doing this.

WYSIWYGS Edtiros Suck, Use WYSIWYM Editors

The basic problem today with developing or using Content Management Systems, Blogs, etc. is that we want to allow the content writers to have a flexible editor with features such as bullets, bold font, different font size, etc. On the other hand, we want to have a strict CSS design for our website, so that the stylesheet will determine how the content will be displayed in a unified manner. With WYSIWYM editors, both goals can be acheived, since WYSIWYM editors generate strict and standard XHTML code, which was designed specifically for this purpose.

Anyways, here is an excellent article about why WYSIWYM editors kick ass:

http://www.456bereastreet.com/archive/200612/forget_wysiwyg_editors_use_wysiwym_instead/

Friday, April 20, 2007

Web 2.0 - Beware!

The new AJAX approach to web design is fun and fascinating, but dangerous at the same time. The main problem with AJAX is that you can't index your site easily. If most of your website content is generated dynamically in the page using AJAX, search engines will NOT be able to index your site content. This issue is supposed to be figured out sometimes, and I'm sure Google is already working on a Javascript / browser simulator to solve this issue out. But Google's solution to dynamic content will never be perfect, because Web 2.0 usually relies on human interaction.

The more concerning issue about Web 2.0 is content stealing - since the basic idea behind AJAX requests is client side data processing (which gives the web much more flexability), the data that is received at the client is plaintext and can usually be parsed in a simple manner (XML or CSV data). The problem is that it becomes very easy to reverse engineer AJAX driven webpages because of the low security implementation. It is much harder to reverse engineer a program and understand how it connects to its remote server, or parse data by yourself from server side web applications. Stealing a webpage written with AJAX can be as simple as copy-pasting functions from the original web page.

So how can these application be protected?

First of all, obfuscation of the data and the code itself. There are program that know how to do it and it might be very helpful to defend against the most common and lamest hackers around. Data obfuscation can be obtained by a simple encryption which is hard to understand and easy to process using Javascript.

The data source itself can be also protected using a referrer check - if the AJAX request came from an unknown page, the service can be blocked. But this can also be easily bypassed by forging the referrer header from the client or from servers that rip the data from the service.

The best technique for protecting AJAX services is using a session - either by using login cookies which the AJAX requests use, or server generated random values that pass back manually from the Javascript itself (the exact same idea, only does not need cookie support and a bit harder to implement). This method is the exact same technique that is used to protect sites from unauthorized users, only that the login sequence is automatic once you enter the main page.

Of course that temporary session cookies are not enough to protect AJAX sites, since another request can be added to extract the session cookie from the main page automatically from the client, which is usually a difficult task to do, exactly as difficult as ripping sites would be, which is exactly what we wanted to achieve.

Thursday, April 19, 2007

SEO?!

SEO (usually rhymes with the word Google) is Search Engine Optimization, the technique (or art) of getting search engines to like your site more.

This is not an article about SEO, but more of food for thought about it. A lot of people make very good money out of the internet, and SEO is a big part of it. If you want people to go into your site and buy / click on ads for you, you need people to know your site exists. Now people usually think that the more you pay Google, the higher you climb on Google's search results. Well, the truth is, Google are good people, therefore they decide whose really the best, and not who pays them the most using several techniques - PageRank and search relevance. That's why many people pay a lot of money for people to do SEO on their sites - because they can't pay Google to do the same. The cool thing about SEO is that it can multiply your revenues by a factor of 10, even more sometimes, just because you show up a few places ahead of what you used to.

Google knows what to find because they're just smart. Smart people know how minds think and act, and that's how the magic is done. Google turns web pages into "mathematical equasions" (well, I can't find a better way to put it), and the pages are indexed using keywords from the page, depending on where the keywords in the page are and what they look like. (For example, keywords in titles have a stronger weight than keywords in small text). But, Google also wants to show you the most relevant pages by giving you the pages they think are better, and that's what makes Google the best search engine around.

Google's method of knowing which pages are better than others (and therefore, more relevant to the search) is the Google PageRank. PageRank is a number from 0 to 10 which means how good and informative your site actually is. Google PageRank is generated once in a while for each page using a lot of variables, such as how original the content of the site is, how often it is updated, how many people visit the page, and many many more techniques.

For example, If someone looks for the word "Flower shop", a page which might be interesting for the person who searched for it should contain the phrase "Flower shop". The more often the phrase is in the page, the more relevant it will be. But if the phrase will show up many times in the page in the same area in the text, it won't count. On the other hand, if the phrase will show up distinctively between very different paragraphs, the phrase will have a stronger signature in the page and will become a better search keyword.

The most commonly known and talked about technique to get a better PageRank on a page is by increasing the number of backlinks to it. The more people link back to your page, the more popular it is, meaning that it will get a higher PageRank. If the sites that refer to your site have a high PageRank, their "vote" will count even more, and your PageRank will be affected more. Using this technique to get you a higher PageRank is very easy (but expensive), just pay people with very well known web sites money so they will put a link to your site from theirs. But, if your site is really good and interesting, people with blogs or sites of their own will link back to your site, and you will get a higher PageRank without actually paying anyone to advertise your site.

Beware, If you do bad things, you get PR0'ed (PageRank 0), which means you'll show up at the bottom of the search.

So what is SEO actually? It's nothing but a technique to make search engines like you more. But eventually, SEO can be legal or illegal. Legal SEO is something that everyone can do. Read a few more articles about search engines, and you'll know how to make your site crawl up the Google ladder. Illegal SEO tricks fool search engines or exploit flaws that allow people to give their page higher PageRank or wider relevance with no good reason. Bottom line is, if you pay guys to do you SEO of that kind, you might get PR0'ed because Google can trace this kind of wierd behaviour.

Finding files over the net

Lately I've been really upset with the fact that finding files over the internet is much harder than it should be. There are billions of files out there just waiting to be downloaded, and no one actually knows where they are...

Now you're probably gonna say "That's bull, you can search for the file on Google and find it easily", but that's where you're wrong. You can only find files which are meant to be found. What does that mean? For example, MP3 files are files which often don't want to be found, meaning that certain search engines (Google, for example) can index them easily but won't do it. On the other hand, you may think that search engines find links to files within pages, which is also wrong because it only indexes the text in the page, and not the anchors themselves. That means that you can only find files over the net by finding pages that lead you to files.

But what if you are looking for a file using a file name instead of its description? What about all the files that are posted in forums all over the world? You'd have to look for the file by entering a description or a file name in a search engine, and then you'd have to look for a link in the page to the file you are looking for (oftenly described using a very big "DOWNLOAD" link). This process may sometimes be annoying, frustrating and time-consuming.

So I got tired of the idea of looking for files over the internet myself. For a start, I wrote a Python based web crawler that searches for a file using a search string, and from there it takes each webpage in the result of the search and read all the links from that page. That gave me a quick and much more comfortable way of looking for files over the internet, without all the fuss behind digging all over the net just to find a small file (which I usually know its exact name, which is just more frustrating).

The script was very useful for downloading specific files which aren't often downloaded. For example: Drivers, DLL's, MP3's, old games, firmware, etc. So I decided to make a webpage out of it: http://www.findthatfile.com. I'm caching my results for webpages that I crawl to speed the search up a bit, so certain download pages which generate download links on the fly (such as download.com) won't work. But I don't really care about these sites since my crawler helps people to find files which are buried in the depths of the internet, and not shareware applications which are already indexed on commercial download sites.

And of course, each file has a link to the page that I found the file in because of copyright issues.

Hello, Blog

Ron: Hello, blog
Blog: Hey, what's up?
Ron: Good.