Affichage des articles dont le libellé est science. Afficher tous les articles
Affichage des articles dont le libellé est science. Afficher tous les articles

15/07/2017

Premier hackathon

Pour la définition du hackathon allez voir sur l'internet mondial encore indépendant.

Premier hackathon donc, un ami me propose de me joindre à lui et à une équipe pour participer à la compétition. 24h dès vendredi soir 18h on est sous la tente de la Startupfest de Montréal.

L'équipe a un projet de biologie en tête avec plein de tuyaux et autres capteurs. L'idée est commencer des cultures d'algues utilisant le CO2 rejeté par des levures dans une bouteilles voisines, le tout relié par des tuyaux flexibles. Suivre les bulles est toujours aussi hypnotisant.

A post shared by jeremie Gerhardt (@mrbonsoir) on

Long story short, je ne suis resté que jusqu'à minuit le vendredi et suis revenu le samedi en fin de matinée. Pendant mon sommeil l'équipe avait bien avancé et les cultures évolué. On n'a pas gagné mais j'ai rencontré des personnages intéressants: un entrepreneur leader de bricobio à ses temps perdus, un biologiste ancien cultivateur de pleurote, un hacker de Québec descendu à vélo à l'occasion, un autre entrepreneur américano français en provenance de Boston et un Panda.

 Une entrée dans la fin de semaine percutante et innovante pour sûr.

25/10/2015

Personal insights on color science

Context
In a previous post I talked about the last event in my field I did attend. Now I want to talk about my perception of this domain which is called color science. I'm pretty sure it can be applied to other fields of research as well.

From the first time I joined this community, from article reader, article contributor to reviewer, committee member and session chair my understanding of what is color science has evolved. One important thing is to stay humble, especially with the new comers. I have been one them, it was impressive. Impressive because you meet the people, authors of research articles that are part of the foundation of you work. You can add a person, a voice to written words, it's actually pretty cool.

There aren't thousand concepts to understand/enter the world of color science. Like in every fields it's about observation and trying to explain what's happening. But here it's all about light - its spectral properties - how we perceive this signal - a single light source to an image in the visible spectrum - and how can we develop robust scientific/engineering "stuffs" around it. What I find interesting is to witness what is the new thing coming each year, how a technical improvement can open a door for further applications.

Color trends
Among the research sub-fields presented at CIC this year I want to come back on four of them.

There is the recurrent discussion about color metrics, from a purely mathematical/geometrical approach to a more perception-wise approach trying to add an average human appreciation of the difference between two signals. Having a good metric is always helpful to evaluate your algorithm/experiment. Over the years the metrics are evolving, context is important (from display calibration to color textile differences...).

There is the what I call "purely geometrical approach" discussion where having a signal as vector of n values - for n wavelength -  a group of sensors - basic configuration made of three basis like RGB basis - you want to know the value of this signal once projected on the known basis/sensors. From that you can jump into optimization, addressing various problems such as finding the scene illuminant/white point, study metamerism. It seems obvious but it's not.

There is printing and 3D printing - there I meant color 3D printing. Just think of how to design a color test-chart for such printing system. HDR display is also coming stronger than ever. What is interesting with these two examples is that they both require to know your workflow, they are the "end" of a process chain: you need to understand the acquisition process to do a good reproduction. Understanding the use of the technology is obviously required.

On the last paragraph one can add the understanding of gamut mapping and how you "move" into your color space as something very important. For printers you have multi-inks system changing the shape of the color space available. For high resolution TV and HDR screen the color gamut shape may not change a lot - almost - but the variability of screen size, intensity scale, technology available make it difficult - to be understood as something cool and challenging for me - to offer a comfortable experience to the user among the different platforms.

Now that I'm a bit more in control with the tools/concepts in my field and sub-fields I have the tendency to prefer the projects combining several concepts - like high quality printing and movie post-production - and I always appreciate to hear how the authors are presenting their projects, which story they are telling us.

23/10/2015

CIC visiting Darmstadt

What is CIC you may ask yourself? It's stand for Color Imaging Conference, a conference about color and imaging. This year it took place in Darmstadt DE. The last 22 editions always took place in the US, last year it was in Boston MA, two years ago in Albuquerque NM, three years ago in Los Angeles CA, four years ago in San Antonio TX and that's it for my involvement. Next stop is San Diego CA in November 2016.

I'm a regular attendee, I joined this community already ten years ago alternating between CGIV, AIC, EI and CIC. Depending of the event you will meet a slightly different crowd or so to say different crowds will meet allowing to go deeper in the various fields represented. But for sure it's about imaging, color, perception, printing, archiving, image acquisition, color management, camera and display calibration, gamut mapping and more.

This year almost 200 persons were attending the event in Darmstadt. There is a kind of routine in such event and being part of the committee allows you to see the people interaction with a special look. It's very special to see the attendees - former colleagues, friends, known members of this community - arriving from everywhere almost - from North America, Europe, Asia, Australia... - and being all jet-lagged. Even if you are traveling in the same time zone you will end up jet-lagged. First of all the schedule is tide and you have to use the "free" time to talk with everybody. Sharing a meal or a beer is usually very appropriate. As a result you barely have time to rest, but the kind of adrenaline you get from meeting the crème de la crème of the color scientists keeps you awake.



30/09/2015

About not being an expert as a data scientist and other tech stuffs

Last evening I did attend a joined Meetup from the Python User Berlin (PUB) and the Zalando Tech Event hosted by Zalando and offering talks on Natural Langage Processing (NLP). Both talks went well and gave two views on the topic: one on the state of the art of the tools for NLP using Python and a second more applied.

The discussions I add after while enjoying a club mate - la boisson des champions - were equally interesting. First of all I started discussing with a expert of NLP trying to explain why I joined this event and what was my link with NLP. In my very recent job experience at EyeEm I just touched the surface of NLP preparing data for Machine Learning (ML) using nltk together with WordNet, ImageNet. Actually I didn't do much of text analysis but batching word definition. In that experiment the text analysis will have come after this step and that's where semantic is jumping into the discussion. Because working with the word dictionary is one side of the problem: you have one word with its definition and often - at least with scientists or engineers - you are in the inverse configuration which is you having words when actually you want to extract a definition, an idea, an information... And I let you google automatic image tagging, deep learning.

After exchanging ideas and experiences about NLP I did continue seeping the offered mate with one Zalando employee. I was curious - as usual - to understand what it means to be a data scientist here. Because if the definition is very general - a data scientist works with data, we are not expert - it's interesting to see how many fields we - I'm one of those people - cover in our daily work. Using the same language - e.g. Python - we can go from signal processing, computer vision, image retrieval, NLP, how to deal with Databases - a year ago I wrote on the topic Databases and natural Langage graphs en stock - how to present your results to non expert by doing nice visualization and many more... So if we are not expert we need to be pretty fast I acquiring skills from various fields and/or use the appropriate tools.

27/05/2015

Struggle for social graph and datavizzz


The holly Grail of the day
Build an interactive data visualization of my own networks where I could jump from one network to the other and navigate in time.  On the paper it sounds easy: use your own network data (facebook (FB), linkedin (LI), twitter, instagram, EyeEm...) to exercise yourself on social graph. In other words use tools from your beloved statistic toolbox (Matlab, Python, R...).

The why
Why, why and why using your own data? First reason and obvious to me, you know the data - or at least part of it - and it should be bit easier to navigate through them. About the first why bother to do that? Once again it's simple and the answer is curiosity. The more people use a buzz word in all conversations the less they understand what it means and I don't like to not understand.

Social graphs are interesting because they illustrate part of our multiple identities - this of course if you decided to look at your own network instead of looking at the interaction between people forming a group which is also interesting (data journalism loves to dissect political social network to find out who are the leaders). We don't know the same people/don't play the same character depending of the network as they describe different interactions (e.g. FB vs LI).

The reverse engineer path
The path I did follow wasn't probably the most efficient but I'm getting better every day. Plotting a social graph isn't the most difficult task. Using gephi you can relatively fast generate beautiful graphs. In parallel I took in statistic and social network analysis to refresh parts of my brain on the topic.

The prototype
As inmaps isn't available any more I ended up on another automatic solution called socilab.con that requires you to log with your linkedin account. It's nicely made, you get a graph and several score values that describe your network and which role you play in it. Sadly it is limited to 500 contacts, so if your contact list is much bigger the analysis is incomplete. But this website allows you to download this version of your contact list. And actually what you are downloading is the formatted data from your LI account under the form of an adjacency matrix. I had to clean a bit the data using Python and Pandas which make any manipulation of csv file a real pleasure.

The adjacency matrix
This matrix - if I understood correctly - should be square where both columns and rows have the same names: your contact name list. Depending of the cell value 0 or 1 you know if your contact know each other or not, the matrix isn't symmetric. It's a particular case of data, because if you look at a FB group of people liking peanut butter toast for diner they may not know each other but they are all connected by their irrational attraction to fatty cream and low safe consideration.

Where the trouble starts
It starts right when you want to access your data... Building by hand this matrix is doable but is a really silly task. And both LI and FB do make the task easy neither. You will need to play with their API (I haven't checked for twitter, instagram and more yet) to access your account and download/build your matrix.



22/05/2015

Deep learning talk @Zalendo Tech Event

First Zalendo Tech Event at their Tech HQ nearby Alexanderplatz yesterday evening. To open their series of Meetup event Zalendo invited Professor Sepp Hochreiter of Johannes Kepler University in Linz to talk about deep learning. 

attentive crowd

About the talk
The talk was good but not adapted to an academic audience. If you are familiar with the topic you probably wouldn't have learned something new. But the talk did lead to interesting - and often expected - questions around and about deep learning. Sadly - to me - it was more where does it work?, what are the best parameters? than how does it work actually? 

As the speaker did remind to us, neural networks (NNs) aren't new on the market. They were discoveries years ago, it was promising and then nothing, other techniques were used, leaving specialists in their niche. I do remember courses during my master in image processing about 15 years ago [in Pierre et Marie Curie Paris VI] where the person teaching and introducing KNN and NNs sounds both excited and disenchanted. This until computers got faster (thanks to cpu, gpu, many-core, cluster, graphic card programming "et j'en passe") and suddenly it was possible to use NNs, to get results, to reproduce them and to beat classification challenges by far comparing to the expert of the field.

For every new promising technique there is the temptation to use if for everything in a brute force manner. But it doesn't work all the time. One remark given by the speaker is these solutions work when you are overloaded with data, when you immersed into data. It's not a surprise that big players such as Google, Facebook, Amazon and more are heavy on growing their deep learning team.

About automation, AI and drugs
You hear and see more and more presentations about deep learning, artificial intelligence (AI) where people are dreaming of AI being able to put words on a given image in a similar way a human will do. It's kind of working but there is no magic. It made me remember about an experiment where the researchers claimed to be able to produce images/video corresponding to the images we see in our dreams. Often people fear - and they can - about computer taking control over us, making decisions for us until we start working for them.

It is interesting to understand why pharmacy companies - those making drugs - are so big into deep learning. Bio-Informatics offer the perfect environment for developing big data solution. Here I'm not talking about the phase where drug need to be tested and evaluated on human but what happen before. Biology and chemistry (or computer chemistry) can be simulated using pretty accurate models, meaning you don't need to run an actual biological or chemical experiment. You can simulate the experiment, generate a huge amount of data and let your algorithm do the analysis. And guess what, computer vision, machine learning, deep learning - not to mention optimization - are part of the solution. And the faster you get your results, the faster you have a new drug to potentially introduce on the market hopefully before your competitor. I'm not sure "normal" people got a glimpse on that side of research, in that field it's actually the biological/chemical experiment that will validate a virtual experiment (remember to watch Terminator 4 or 5 at leas the last on screen...).

About the big brain project and graphic cards and evolution
Research is cool. It's very interesting to see how connections/links between highly specialized fields are happening to build a new framework for research. The big brain project (not sure about the name but there is the US and the EU version) is the perfect example, different fields from neuroscientists to computer graphics and hardware manufacturers need to collaborate to build this virtual brain model.

One of the last comment from the speaker yesterday had a pertinent echo in my head. This comment illustrates perfectly how technology is evolving and frameworks are crossing their paths. He told us that graphic card manufacturer (such as nvidia to not name them) are now developing hardware dedicated to run deep learning process, once again the hardware architecture helping to fasten a programmed algorithm. But until when and is it a good approach? 

Years ago and not so long time ago when computers were already getting faster, people were designing hardware to run image processing/computer vision algorithms. This because the computers in their at-this-time state weren't fast enough. Like the brain was too small and needed to grow or modify its physical body to evolve. But then computer became faster and those special design weren't adaptable enough, too specialized. I feel that we are living a similar state with deep learning. The question will be is hyper-specialization of computer hardware the solution - momentarily for sure - for deep learning or not?

About the future
We are all doomed. Soon computer will be smart enough to redesign their body when they will reach their limits to overpass them. I haven't any spoiler about how and when, out Mayan friends had a big fail about it three years ago, we have to be patient.

  

09/04/2015

Deep learning (ou deep learning in French)

What was your question already?
How to explain deep learning to your friends, family members, neighbors, random stranger, dog? A very good question indeed. Rather than going deeply into neural networks and other festivities let's start with describing the problem(s) we want to solve. Or least let's give an example of what we are trying to do here.

Over the years I had to come with strategies if I wanted to explain what I do for living. Giving keywords such "color science", "computer vision", "image processing", "digital photography" is usually not enough or saying "I do work with images" neither. I always found interesting to answer the question "why to you want to do that?" or "which problem do you want to solve?". So to explain what I can do I try to give an idea of the tasks I have to solve.

What is the problem you are trying to solve already?
In some way asking these questions is already machine learning/deep learning-ish approach of solving a problem. In theory if someone asks you to solve a problem he knows the kind of results he want to obtain for a given input or starting point. What he doesn't know is what is happening between these two stages. Applied mathematics and optimization are a reasonable standard solution: you develop of model that recreate more-less accurately what is happening between these two stages, then for a new entree point your model will predict what an output will be.

I'm sure "big data" is an expression you have heard in the past years or months. It has of course different meaning depending who to is giving a definition. But, coming back to images and the incredible amount of images we are producing daily there is a need to develop solutions, tools to be able to interact with these images. You have in your hand an extremely large image database and using keyword as a search query isn't enough anymore. So here is the problem: how to navigate, how to browse into large image database in a more natural way? There is a bit of database here but that is not the main point of my article, check my past post on graph and database if you are interested.

Face recognition to recognition of everything
Working with images is fascinating, you see one image and automatically you extract some of its  information. Of course there is a long learning curve, when you see a tree, a car, a known object in a picture you don't even realize it, you know, you have learned over the years you spent on earth to recognize, categorize, organize the continuous stream of visual information that come to your eyes and is later processed in your brain.

If you think of face recognition, the mathematical tools are now pretty standard. We can with high probability find out faces in images, classification comes after the recognition. And if you train your model you will be able to recognize semi automatically in a database faces of different persons as the tools/filters can be tuned for a given target. It can be scary of course if the threshold that decide for a true recognition/classification isn't verified by a real human and that action lead to a rocket launch. Actually any automatic action issued from an algorithm decision having impact on a human being is pretty bad (hello mass surveillance and hello Terminator). You want help from robots not to help robots or it's too late anyway.

An idea behind deep learning is to be able to learn what are into images - in a similar way as we human do - to extract features and to perform tasks on other images based on a trained neural network. I'm making shortcuts but that's the idea. To understand and to later mimic how information is circulating into the brain has been a dream of many researchers. Neural networks go into that direction. If a few years ago the algorithms were limited because of computer power the global picture is different now.

What is also interesting is that new strategies had to be developed to overcome the overload of data. In a way the system were "over learning" and people talked about over-fitting the data. And it makes sens. If I'm not too mistaken our brain is not indefinitely expandable, meaning we are sorting information continuously. One big part of these tools is to perform drop-out which can be explained as "now that our system can learn we have to teach him to forget part of what he knows in real time".

Cross disciplines 
A chance I see - for me - is the need in some industries for expert being not only expert in one field. Specially for this kind of large scale problems involving images, computer vision, real time and fancy applied research projects. To know only about machine learning or statistic is not enough, to know both about computer and machine learning tools is better.

[We talk later about existing and possible applications.]