Hello folks,
After a long time, something actually technical coming up here. Actually after hearing about GP (Guruprasad, my senior)'s project (again) from Salai. I and him, have decided to put up a small working prototype. Working towards the same, i came across an interesting problem suitable even for my RE iBlue - Semantic Analysis.
For my KBEngine implementation, I downloaded structured content from Wikipedia in readymade RDFs and used data store like Jena TDB to store, index and retrieve the contents upon depend. Now, its time I start structuring my own data from any data source.
Link: http://ashwanthkumar.in/labs/iblue/
Above link, shows you a pre-alpha working prototype attempt in doing so. It uses various available webservices for structuring of data, mainly using Apache UIMA for processing. The hosted version uses external content to process, while the actual RE implements the same using Apache UIMA in its own infrastructure.
Link: http://ashwanthkumar.in/labs/iblue/iblue.php
Above link shows you a sample of the analysis format using this blog post content as its source.
Have a look at it, and please do comment your feedback.
PS: Well I'm, thinking of submitting this concept with my RE for Daksh 5.0 Technovation.
Showing posts with label tgmc. Show all posts
Showing posts with label tgmc. Show all posts
Thursday, February 17, 2011
Tuesday, December 14, 2010
Social Networking Platform (IndiKonn)
Project Name - IndiKonn
Project Scenario - Social Networking Platform
Team Members - Lakshmi Narayanan B, Divya K, Jayalakshmi S, and Prasanna Kumaor
You can also download it from here - http://goo.gl/W2c8j
Project Scenario - Social Networking Platform
Team Members - Lakshmi Narayanan B, Divya K, Jayalakshmi S, and Prasanna Kumaor
You can also download it from here - http://goo.gl/W2c8j
Sunday, November 14, 2010
Research Engine - Alpha (under way)
Research Engine Alpha, is under construction and will be out for testing soon. Its primarily built on Apache Solr server. Its plugin architecture helps a lot to include a custom components performing the required job, and still remaining to be a stable search server.
I thought it would be easy to start with Solr, and go as required.
ETA: 45 days
- Ashwanth Kumar
This is an update post, used as a log for me and my teammates working on the project to record the happenings of our work. This is a personal effort. If you plan to use this intelligent property, please drop a message at: http://goo.gl/nqkWs
Sunday, October 17, 2010
Research Engine - A new identity
Well, this is yet another series on my TGMC project this year, "Research Engine". I've been thinking off late so much about Research engine, especially after my break up with my girl few weeks ago.
- Ashwanth Kumar
"If i would make it a search engine, it would still like be one of them. May be one with a better filtering capability of all others, but there is no innovation in that. Search is still not a solved problem! I still believe in that, but what can Research Engine can possibly achieve?!", said me to myself.
I was pondering over this for a long time, until now. I again came up with a new improved version of Research Engine, one already but not completely visualized by my team mate, Salaikumar @ Saravanan (http://goo.gl/tfrx).
We came up with another way of representing information to the users. All search engine, gives you pages and pages of links, to some so-called relevant information for your search query. We though why not give the users directly the information what they are looking for, instead of giving them links to go and search from.
So we decided to have a user view layer like that. It is information centered (re)search, so that you don't have to re-search again else where.
Well you should be thinking, "Isn't this what Wikipedia does in the first place?". Actually the answer is, "Yes!". You can also call Research Engine to be the next encyclopedia, but the potential of RE is more than just displaying information.
Major features or improvements over a wiki is that,
- All the data collection process is completely automated.
- All the information you see is live data (possibly some milliseconds old)
- All the information you get is presented after analysis of same content over a variety of authenticated websites.
That is with this post as of now. If you feel, RE can have more features or if you have a suggestion for the same, please drop your views on the issue, as comments. As i said, i want this to be a community driven Web 3.0 Technology, that makes the lives of the people easier.
- Ashwanth Kumar
Friday, September 3, 2010
Research Engine - Project Proposal Accepted
Horray... Horray.. Our project proposal, for TGMC 2010, Research Engine has been accepted by the TGMC 2010. Now, its time to kickstart our project officially in full fledge.
Hell Yeah!!
PS: This is an update post.
Tuesday, August 24, 2010
Our TGMC Group - Tech Buddies
Created our TGMC Group - Tech Buddies (http://bit.ly/bj32sz), as per the TGMC rules. Except that i forgot to register first! (:0)
All other information can be found there.
PS: This is just an update post.
All other information can be found there.
PS: This is just an update post.
Friday, August 20, 2010
TGMC as a Final Year Project?!

Guys,
I was browsing through the wikis' of TGMC '10, and i harped upon this (http://bit.ly/d3a0sp).
Wiki Page titled, "TGMC Project as Final Year Project", For all my final year friends, do u think will be able to implement this in our university or the campus?
Wiki Page titled, "TGMC Project as Final Year Project", For all my final year friends, do u think will be able to implement this in our university or the campus?
Please do comment, what do you think about this!
Wednesday, August 18, 2010
Components of Research Engine - Developer Preview
Okie, this is going to be the first component proposal for Research Engine, by Ashwanth (myself) and Kirubaharan.
Any comments, and feedback regarding this is welcome.
Any comments, and feedback regarding this is welcome.
Ya, I know, the project is becoming more and more community based. I dont mind, if there is a fork too. Just let me know there is one! I would be happy to know it. Once we start Coding, CodeBase is also planned to be publicly available too.
TGMC '10 Project Scenario
TGMC 2010, now has some new rules regarding custom Project Scenario being sent to him for approval, and then development of the same.
Source: http://bit.ly/di8WDW
Well, Research Engine, does need approval i think?! Anyways, I'm currently working on the same against the TGMC Scenario Template. I'll post the Link, once i'm done with it.
UPDATE on 20th August, 2010: Our TGMC Project Scenario (http://bit.ly/bdETEv). We've send this for approval, and awaiting for their response.
PS: This is just an update post. To log the status of the app.
Sunday, August 15, 2010
TGMC 2010 is Out!

TGMC 2010, is out with a bang. This time, it is expected to be with more fun and learning experiences.
The site was quoted as: "The Great Mind Challenge (TGMC) is back in a new avatar!After a superb year in 2009 where we saw more than 100,000 students participating with some excellent projects, this year we’re proud to announce the launch of TGMC 2010.This year, there is unprecedented focus on the most important aspect of TGMC. You.As a student or a faculty member, TGMC 2010 is the forum for you to come forward, take your destiny by your hands and make it happen. And we put it in the form of this simple mantra.
‘Initiate. Collaborate. Innovate.’YOU need to initiate the chain of events that will take you to the brink of Success.
YOU need to collaborate, with your peers, your faculty and IBM to ensure you achieve that Success.
And YOU need to innovate to ensure that together, you achieve not only Success, but a lasting place in the Halls of Fame.Technology is the vehicle. YOU are in the Driver’s seat. And it’s going to be a great ride.
Saturday, August 7, 2010
Crawl + Index Module Test - Successful
Today, after a long sleepless night, crawl and indexer search module of Nutch has been tested, and this blog has been taken as the test benchmark. Results seems promising.
Tech Info: Test was performed on a single system, running the following config: 2 GB RAM, Ubuntu 10.04 Desktop, 160 GB HDD. Time taken was ~75 secs @ 2Mbps connection.
PS: This is an update post.
Tech Info: Test was performed on a single system, running the following config: 2 GB RAM, Ubuntu 10.04 Desktop, 160 GB HDD. Time taken was ~75 secs @ 2Mbps connection.
PS: This is an update post.
Research Engine - Code Name: iBlue
Now, its official (from the team + mentor), the Research Engine has been code named: iBlue.
Logo coming soon.
PS: This is just an update post.
Logo coming soon.
PS: This is just an update post.
Wednesday, July 28, 2010
Hadoop Cluster Deployment + Step-By-Step Process
I've successfully deployed a small cluster of 3 nodes on Hadoop platform. I mark this as the 1st success towards the long road for Research Engine. It took me a while to understand the basics (since this is the 1st time) but it was such a wonderful experience.
The Cluster Specs're:
Tests: Grep for Map/Reduce, Content Duplication by going on a copy of 500 MB replicated over 2 nodes.
Info: 1 NameNode, 1 DataNode, and 1 JobTracker
Below is the Step-by-Step procedure to deploy a Hadoop Cluster (for Learning purposes only. This can't be used as such in production environment. Please refer to Official Docs, and latest release for that). See the disclaimer on the bottom before even you start reading beyond this.
Disclaimer: This is for my future reference. I don't take any responsibility over physical/mental/any other type of damage that may arise on following the above said process.
The Cluster Specs're:
- Core - Ubuntu 10.04 - 2 GB RAM
- Core - Ubuntu 9.04 - 1 GB RAM
- Virtual PC (VBox 3.2.4) - Ubuntu 10.04 - 512 MB RAM. (Host is 1 machine)
Tests: Grep for Map/Reduce, Content Duplication by going on a copy of 500 MB replicated over 2 nodes.
Info: 1 NameNode, 1 DataNode, and 1 JobTracker
Below is the Step-by-Step procedure to deploy a Hadoop Cluster (for Learning purposes only. This can't be used as such in production environment. Please refer to Official Docs, and latest release for that). See the disclaimer on the bottom before even you start reading beyond this.
- In this steps, i shall assume you've 3 -4 systems, on a network and each of them running on Ubuntu 9.04+ with sun-java6-jdk and ssh packages installed. Its preferable to use a new system installation, though its not mandatory.
- Due to some issues with Hadoop 0.20.* (latest stable as of writing this post) we shall now (currently) use Hadoop 0.19.2 (stable). You can get a copy of yours from: http://apache.imghat.com/hadoop/core/hadoop-0.19.2/hadoop-0.19.2.tar.gz (53 MB).
- Create a new user for Hadoop work. This step is optional. Its recommened, as the path HADOOP_HOME is the same in the cluster.
- Extract the hadoop distribution on your home folder (u can extract it anywhere though). So, your HADOOP_HOME will be like: /home/yourname/hadoop-0.19.2
- Now, repeat the steps in all the nodes. (make sure the HADOOP_HOME) is the same on all the nodes.
- We need the IPs of all the 3 nodes. Let them be: 192.168.1.5, 192.168.1.6, 192.168.1.7. Where *.1.5 is the NameNode, *.1.6 is the JobTracker, these 2 are the main exclusive servers. You can find more info regarding them here (http://hadoop.apache.org/common/docs/r0.19.2/cluster_setup.html#Installation). Node *.1.7 is the DataNode, which is used for both Task Tracking and storing Data.
- U'll find a file called: "hadoop-site.xml" under the conf directory of the Hadoop distribution. Copy and paste the following contents between <configuration> </configuration>
<property>
<name>fs.default.name</name>
<!-- IP Of the NameNode -->
<value>hdfs://192.168.1.5:9090</value>
<description></description>
</property>
<property>
<name>mapred.job.tracker</name>
<!-- IP of the JobTracker -->
<value>192.168.1.6:9050</value>
<description></description>
</property><property>
</property> - Make sure the same is done for all the nodes in the system.
- Now, to create the Slaves for the NameNode to replicate the data. Go the HADOOP_HOME directory in the NameNode. Under the folder "conf" you should see a file called slaves.
- Upon opening slaves, you should see a line with "localhost". Add the IPs of all the DataNodes you wish to connect to the cluster here, one per line. Sample Slaves will be as follows:
localhost
192.168.1.7 - Now, its time to kick-start our cluster.
- Open terminal in the NameNode, go to HADOOP_HOME.
- Execute the following commands:
# Format the HDFS in the namenode
$ bin/hadoop namenode -format
# Start the Distributed File System service on the NameNode, which will ask you the passwords for itself and All the slaves, to connect via SSH
$ bin/start-dfs.sh - Your NameNode should start and be running. To check the nodes connected to your cluster, go to step 19 and come back.
- Now, its the JobTracker Node. Execute the following commands:
# Start the Map/Reduce service on the JobTracker
$ bin/start-mapred.sh - The same process follows for JobTracker. It asks all the password for itself and all its slaves (did i tell u, u can also add slaves to JobTracker?. Its the same process as NameNode, just add the IPs to the slaves file of JobTracker Node's Hadoop distribution), to start the Map/Reduce service.
- Now, that we're done starting the cluster. Its time to check it out!
- In the NameNode execute the following command:
# Copy a folder (conf) to HDFS - For sample purpose
$ bin/hadoop fs -put conf input - If you go to http://192.168.1.5:50070, on your browser. U should see the Hadoop HDFS Admin interface. Its a simple interface created to meet the purpose. It shows you the Cluster Summary, Live and Dead Nodes etc.
- U can browse the HDFS using Browse the filesystem link on the top-left corner.
- Go, to http://192.168.1.6:50030, to view the Hadoop Map/Reduce Admin Interface. It displays the current jobs, finished jobs etc.
- Now, its time to check the Map/Reduce Process. Execute the following:
# Default example code comes along with the distro.
$ bin/hadoop jar hadoop-*-examples.jar grep conf output 'dfs[a-z.]+'
- http://hadoop.apache.org/common/docs/r0.19.2/quickstart.html
- http://hadoop.apache.org/common/docs/r0.19.2/cluster_setup.html
- http://hadoop.apache.org/common/docs/r0.19.2/mapred_tutorial.html
Disclaimer: This is for my future reference. I don't take any responsibility over physical/mental/any other type of damage that may arise on following the above said process.
Thursday, July 22, 2010
Which Data Warehouse Infrastructure to use?
OMG! Problem of selection of components has finally started again. Question is: Which data warehouse infrastructure to use for Research Engine?
Available choices are:
Available choices are:
- Apache Hive - http://hadoop.apache.org/hive/
- IBM InfoSphere Warehouse - http://www-01.ibm.com/software/data/infosphere/warehouse/
- Mike 2 - http://mike2.openmethodology.org/
- MySQL (Really ?) - http://opensourceanalytics.com/2005/11/03/data-warehousing-with-mysql/
Update: Since its TGMC, i'm sticking with InfoSphere.
Research Engine - Work Flow
A proposal for the event work flow for Research Engine, If u've suggestions or improvements, please do post it as a comment.


- Get the user input search terms or Query
- Find the Model (or Domain) at which the Query belongs to. This step is to find the model of the user query using keywords. The purpose of this step is to identify as much related models of the query as possible, for the Query processing based on Language Processing (LP) techniques.
- Once the list of related models is identified, the query is now under the process of Language Processing (LP). This step ensures the evolution of the Research Engine, over time. This step does the following work: Understand the Query, Identify the exact related models (if any) or Create new Models (if none).
- Once we identify the Models regarding the query, query the Model Data Store (DS) to fetch the related information about the model (subset of the model).
- The output of the previous step gives all the related information, the user wants. Now, all that is left is to output the processed info in any format of choice (depending upon the application).
Wednesday, July 21, 2010
Research Engine - Interactive Search Engine
Atlast our mentor accepted our proposal, "Research Engine". Its basically an incremented version of Semantic Search Engine. Its being planned to be built from down to top in a pure interactive way. Since, too much of interaction makes the user lazy or dislike the concept, we're planning to extract meta data from social networking profiles of the users(like Facebook, Orkut, Twitter, MySpace, etc.) to automate the process of interaction and improve it, dynamically
Also, we're planning to build this project on-top of Nutch. Many modifications are required to make it a semantic search engine. DB2 9.5 Enterprise, Jena, WASCE, Hadoop, are some the major components to be included in the project.
PS: I'll try to make updates like this regular, but not sure about it either.
Also, we're planning to build this project on-top of Nutch. Many modifications are required to make it a semantic search engine. DB2 9.5 Enterprise, Jena, WASCE, Hadoop, are some the major components to be included in the project.
PS: I'll try to make updates like this regular, but not sure about it either.
Tuesday, July 20, 2010
My TGMC teammates this year
Atlast after a long struggle for team members with vibrant interest and enthu to match my frequency, i got myself the best pieces of SRC. Below are the names in Alphabetical order:
- Ashwanth Kumar - III CSE
- Kirubaharan A - III CSE
- Saravana Kumar - II CSE
- Swetha S - II CSE
Subscribe to:
Posts (Atom)