Showing posts with label semantics. Show all posts
Showing posts with label semantics. Show all posts

Tuesday, November 15, 2011

Image processing meets Semantic Web

For the past 2 days, I was working along side with a friend of mine for her final year project topic title. Her guide restricted the domain only to "Image Processing", for she had something to do with it I guess. Anyways after browsing and downloading about 40+ papers against Image processing. I really liked this one. 

Yesterday we were preparing the abstract and after a lot of local misunderstandings, the guide finally signed her abstract. Today morning, when I was preparing for my semester exam, this thought struck me and I thought before I forget, let me document it. So here it goes.

Image Processing + Semantic Web => My Kind of Vision

If you took time to read the abstract in the above link you would know how wonderfully they had brought about a new data structure for image annotation. Their motive is to build an online collaborative image annotation tool, something like LabelMe. The main feature being - modular design,  ability to import other online (LabelMe, Flickr) and offline (Caltech101, Lotus Hill) datasets into the system. Apart from the text based human annotation, it can also embed the low level image details like color histograms, etc. (I'm relatively new to Image Processing and let me just skip the rest of the details with the thought not to confuse you).

Well, this is already there and she has decided to implement such a tool with other added features like Web Services to search through the annotations, and more. Rest of blah.. blah.. content goes here.

When I was browsing through the dataset of LabelMe, I figured out something. Its like the usual temptation to study well only outside the exam hall. This is what I concluded myself with before I ran out of time for my E-Commerce exam.

I've a collection of images that were annotated very well (atleast decently well) to make it more human friendly and make the images more semantic. That does not make it Machine friendly does it? (or am I being carried away?) 

Another MIT Media Lab project is ConceptNet5. This provides general usage common sense knowledge to the computers in my most favorite language - JSON. It contains around 15+ million entries in it. 

So, this is what I'm going to build (hoping my guide would approve it). 

  1. Taking the existing annotated image dataset, apply ConceptNet common sense knowledge (on the annotation of the images) to make the images more semantic and machine friendly. Now at this point machine can learn from the annotated image - Let us call it TEACH mode
  2. Build an index of properties with all the available images to enable image or object recognition via a Search Interface - Let us call it SEARCH mode
  3. Build a reasoner on top of the ConceptNet relations and annotation currently processed. This helps make conclusions based on the image annotations - Let us call it PROCESS mode
Possible Applications 
  1. Now, combined the power of Relations and Reasoner I would be able to fetch dynamic content (from Google and Wikipedia) about a particular annotated object within a image when queried. 
  2. With the power of Image index thus created, I can recognize the object and thus automate the process of annotation. thus providing the previous application in a more automated way. 
Things I have to learn and random notes
  1. Basic image properties that I would be indexing with the images
  2. May be use SVG or a another custom data structure to store the index
  3. Bridge the relations and concepts to Predicate logic reasoning
  4. Modify a Machine Learning algorithm (something like Naive bayesian or even more sophisticated to make the image learning possible)

If anyone of you find any more features worthwhile to be added into the system, please feel free to post it as a comment. Probably this Idea isn't new at all. I didn't take time to Google about it. Let me know if its already there, probably we can build something really even more useful on top of that. 

Wednesday, May 4, 2011

iBlue Semantic Extractor - Now Speaks 52 Languages

Another Good morning to you people. Today, I've an interesting update regarding iBlue Semantic Extractor. Semantic Extractor component was initially build as a project prototype for Mr. Guruprasad S. Srivatsav back last year. It simply extracted entities out of the text, find their relations, and categories. When it was first created, it had a 1000 word limit, and English language only detection.

Today, I present you iBlue Semantic Extractor with 52 language support. You can enter the text in any of the 52 languages, and it is translated to English dynamically and then identifies the entities, relations and categories.

52 Languages supported are: Afrikaans, Albanian, Arabic, Belarusian, Bulgarian, Catalan, Chinese Simplified, Chinese Traditional, Croatian, Czech, Danish, Dutch, English, Estonian, Filipino, Finnish, French, Galician, German, Greek, Haitian Creole, Hebrew, Hindi, Hungarian, Icelandic, Indonesian, Irish, Italian, Japanese, Latvian, Lithuanian, Macedonian, Malay, Maltese, Norwegian, Persian, Polish, Portuguese, Romanian, Russian, Serbian, Slovak, Slovenian, Spanish, Swahili, Swedish, Thai, Turkish, Ukrainian, Vietnamese, Welsh, Yiddish

You can visit iBlue Semantic Extractor Component for free preview here - http://ashwanthkumar.in/labs/iblue/

Next Milestone - Identify the entities natively from the given language.

Tuesday, April 26, 2011

Sentiment Analysis - iBlue Component Preview

iBlue is fast evolving from a dream to a reality. After working on Semantic Web in depth for more than a week, I learnt many principles, theorems (I didn't even bother to study theorms' for my Maths papers), standards, and many more.

You might be had a look at my Spam Detection plugin for Elgg, here. It was during working for this, I came to know about Bayesian Filters and its usage in SPAM Detection, text classification, etc. It is one of fundamental machine learning techniques. Blah.. Blah.. You can find more details at the wikipedia - http://en.wikipedia.org/wiki/Recursive_Bayesian_estimation

Thats it with the introduction, and now I welcome you all to test drive my very first Bayesian filter implementation - Sentiment Analysis. It is every similar to the working of Semantic Extractor, except it gives only one information about the text.

Is the given text a positive feedback or negative?

It returns the final percentage of both the cases. So, just go ahead and give it a try.

Technical Specs for Nerds - The filter was trained using the test data from the dataset provided by Mark Dredze in their Multi-Domain Sentiment Dataset. I took around 25000 Amazon reviews from the dataset to train the filter, from multiple product categories.

Demo for : http://ashwanthkumar.in/labs/sentiment/sentiment.php - It uses the content from Mashable for Samsung Galaxy S Android Smart phone (Link: http://mashable.com/2010/07/26/galaxy-s-review/)

Any feedback is highly appreciated.

Saturday, April 23, 2011

Semantic Week

If my entire last week was with Elgg, this entire week was with Semantics. Semantics in various aspects, you should have already seen my Semantic Social Viewer (http://ashwanthkumar.in/labs/facebook/), did some really super stuff with Guruprasad S. Srivatsav for his final year project.

I've been looking around WWW for some Social Computing stuff, and things. Found a lot of new ideas, to work on this summer.

Most important of all my GSoC results are coming out on 25th of this month. As, I've already told, I've no hope for that at all. I got introduced to a guy, whom I check'd out to only find, that he has applied to the same Organization, and same topic as mine, with 150% more sophisticated profile. So, the dream of $5000 to fix my PC is all lost! Still, #elgg is always my best friend. I'll continue to contribute to Elgg always. I'm waiting the results to release my plugin.

So, nothing more interesting yet! Got my semester practicals coming up. Yeah, its the same procedure of Ctrl + C, & Ctrl + V and some 30 min vetti time waste. Pracs all over. Theory, is all ready to upfront me, while I ain't.

One more to go after this, count down started! =D