Alphabet Recognition Using Hand Gestures – A Deep Learning and Computer Vision Project

Machine Learning

Data Structures

Algorithms

Problem Solving

Home >

Machine Learning >

Alphabet Recognition Using Hand Gestures – A Deep Learning and Computer Vision Project

Post author: Shivansh Joshi

Post published: February 25, 2021

Post category: Computer Vision / Machine Learning / Projects

Post comments: 1 Comment

This is a tutorial on how to build a deep learning application that can recognize the alphabet written by an object-of-interest ( red colour object) in real-time. Also visualizing the alphabet on blackboard with contour for debugging purpose.

OCR model is created on the top of Keras API for this project, as we have to predict the alphabet in real-time.

This deep learning application in python recognizes the alphabet through gestures captured real-time on a webcam. The user is allowed to write the alphabet on the screen using an object-of-interest like a pen or something similar.

Here object-of-interest is referred to as a particular colour in our input frame like in our case it’s a red colour pen. You may change the colour according to your will by changing the redLower and redUpper .

Note- There must not be any other red colored object in our input frame. As it might cause disturbance for our application .

You can access the full project code:

https://github.com/hellomlorg/Alphabet-Recognition-Using-Hand-Gestures

The code is in Python version 3.7, uses OpenCV , TensorFlow and Keras libraries.

The “Extended Hello World” of object recognition for machine learning and deep learning is the EMNIST dataset for handwritten letters recognition. It is an extended version of the MNIST dataset. Like we program Hello world in C++.

In this project, we have directly imported the EMNIST using mnist.loader. It comprises a total of 1,24,800 images of A-Z (randomly stored).

Each of the letters in Emnist is stored as a numbered array(28 x 28) as shown below.

1.1 Loading the Dataset :

We use Python’s mnist library to load the data. This dataset consists of 124800 images of handwritten alphabets. Each image is in the form of an array of size 28 x 28. i.e each image are 28×28 pixels each.

Lets now get the data ready to be fed to the model. Splitting the data into train and test sets, standardizing the images and other preliminary stuff. Splitting the data is done using the predefined function of Scikit learn i.e the train_test_split which split the input into two .ie. the training set and the testing set. Using the test ratio that I have set to 0.2 in my case , that means 20 % of 124800 images will move to testing set and other would go in training set , which we would use for training the Keras model.

We also apply One hot encoding, i.e a very important step, for training the model using Keras or Pytorch. Ex- like if our label is ‘C’ .i.e 2 so it will make an array of 26 size with all its element set to 0 except that at index 2.

1.2 Defining the Model :

In Keras, models are defined as a sequence of layers. We first initialize a ‘Sequential Model’ and then we add the layers with respective neurons in them.

In our case a Dense layer is added to form an inner layer with 512 nodes and an activation function of “Relu” ( Relu has a straight line y=x graph for all input x>0 and a y=0 for all x<0 ) is used. You may use any other activation function, but it works well in my case.

The model, as expected, takes 28 x 28 pixels (we flatten out the image and pass each of the pixels in a 1-D vector) as an input. The output of the model has to be a decision on one of the letters, so we set the output layer with 26 neurons (the decision is made in probabilities like A class has the probability of 0.8 and B is predicted as 0.2). Finally, a softmax layer is added to find the class of the alphabet i.e A-Z matches to 0-25 respectively.

When working with image data we have to distinguish how we want to encode it. Since Keras is a high-level library that can work on multiple “backends” such as, TensorFlow, Theano or CNTK, we have to first find out how our backend encodes the data. It can either be encoded in a “channels first” or in a “channels last” way which is the default in Tensorflow in the default Keras Backend. So in our case, when we use Tensorflow it would be a tensor of (batch_size, rows, cols, channels). So we first input the batch_size, then the 28 rows of the image, then the 28 columns of the image and then a 1 for the number of channels since we have image data that is grey-scale.

1.3 Compile the model:

Now that the model is defined, we can compile it. Compiling the model uses efficient numerical libraries such as TensorFlow or Theano. Here, I specify some properties needed to train the network. By training, I am trying to find the best set of weights to make the decision on the input.

At the so-called Backend, backpropagation took place to assign the correct weight to each and every node in the layers. In each epoch, the error is calculated and managed using the Optimizers like – Gradient Descent, Adam Optimizer, etc. In our case, I have used Adam Optimizer as it works well in the case of categorical classification. I must specify the loss function to use to evaluate a set of weights, the optimizer used to search through different weights for the network and any optional metrics we would like to collect and report during training.

1.4 Fit model:

Here, I train the model using a model checkpoint, which will help us save the best model (best in terms of the metric we defined in the previous step). The best model will be finally stored in my_best_model.h5. It would help us save the best model with the highest accuracy and lowest loss among all the model that is passed during each Epoch.

1.5 Plotting the graph and evaluating the accuracy:

I plot the graph using the matplotlib .i.e the loss graph and accuracy graph. On seeing the graph it’s noticed that the loss is decreasing as we move forward on epoch while the accuracy increases meanwhile. The “TensorBoard” can also be used for plotting, it plots in realtime while training the model. Finally, I find the accuracy of my model i.e 85 %. (much better than Before !! ).

Finally, there is no need to save the model using pickle as it’s already saved during the training with the name my_best_model.h5. Now, Its time to implement that model using a webcam with a real-time gesture. That is if we hold a red coloured object in front of the camera it will detect that and if we make any alphabet shape using that, it will show the prediction and will also speak the answer.

2.1 Describing important variable :

Before we look into the recognition code, lets see important variable .

First, I load the models built in the previous steps and storing them in the model. I then create a letters dictionary, redLower and redUpper boundaries to detect the red coloured object, a kernel( dilation, opening and erosion ) to smooth things along the way, an empty blackboard to store the writings in white (just like the alphabet in the EMNIST dataset), a deque to store all the points generated by the object-of-interest, and a some of the default value variables.

2.2 Capturing the writing :

Once I start reading the input video frame by frame, we try to find the object-of-interest. We use OpenCV’s cv2.VideoCapture() method to read the video, frame by frame (using an infinite while loop), either from a video file or from a webcam in real-time.

In this case, we pass 0 to the method to read from a webcam , because we want frame input without any delay. Once we start reading the webcam feed, we constantly look for a red colour object in the frames with the help of the cv2.inRange() method and use the redUpper and redLower variables initialized beforehand. Once we find the contour, we do a series of image operations and make it smooth.

Smoothing just makes our lives easier. It takes care of the image that we pass to the input OCR model is fine without any noise in it. Once we find the contour (the if condition passes when a contour is found), we use the centre of the contour (red colour object) to draw on the screen as it moves. The following code does the same.

The above code checks if a contour is found and if yes, it takes the largest one, draws a circle around it using the cv2.minEnclosingCircle() and cv2.circle() methods, gets the center of the contour found with the help of the cv2.moments() method. In the end, the centre is stored in a deque called points so that we can join them all to form full writing.

We display the drawing on both the frame and blackboard . One for external display and the other to pass it to the model.

2.3 Cropping the writing and passing it to model:

Once the user finishes writing, we take the points we stored earlier in deque, join them up, put them on a blackboard and pass it to the models. The control enters this elif block when we stop writing (because there were no contours detected as there was no red coloured object in the frame ). Once we verify that the points deque is not empty, we are now sure that the writing is done. Now we take the blackboard image, do a quick contour search again (to crop the alphabet out from that blackboard ). Once found, we cut it appropriately, resize it meet the input dimension requirements of the models we built i.e., 28 x 28 pixels. And pass it to the OCR model that we have created earlier.

2.3 Finally Showing model prediction on the Original Frame :

We then show the predictions made by our models on the frame window. And then display it using the cv2.imshow() method and also dictate the result using the library pyttsx3. After falling out of the while loop we entered to read data from the webcam.

On pressing ENTER , we release the camera and destroy all the windows.

In this tutorial we created a deep learning model to predict the hand written alphabet and implementing it using the computer vision concept of OpenCV libraries . Using this Project basics , we can also implement the project of Identifying digit using finger count , Soduko Solver using Deep Learning , RealTime video to pdf capture and Other Computer Vision Projects.

Hope you enjoyed this article. For more such amazing articles, check out helloml.org . If you want to improve this article or report something incorrect, please send a mail to [email protected] .

Any copyright/wrong information claims will be directed to the author/s.

Computer Vision enthusiast || Deep Learning || Django and web development

Tweet

WhatsApp

Telegram

Register

Lost your password?

Introduction to Binary Search Tree

Error Analysis in Machine Learning

A Very Quick Intro To Hive

Logistic Regression as a Neural Network

Iterative Dichotomiser 3 (ID 3)

Join our internship program to learn and grow. All with a passion for technology can join.