Custom NLP Library

This NLP class handles file-based text processing, including reading, tokenizing, and creating a sorted bag of words. It supports detailed cleaning steps to normalize contractions, hyphens, and punctuation for accurate analysis.

The Lemmatize class applies a finite state transducer (FST) model to systematically break down words and reduce them to their lemmas by following character transitions. It uses a trie structure to store transitions between characters and determine if a word matches a stored lemma, enhancing accuracy by filtering out proper nouns through the NLP.tokenizer.

Dependencies

numpy

Install Dependencies

pip install numpy

Run Tests

Clone the Repository and run:

python run_tests.py

Name		Name	Last commit message	Last commit date
Latest commit History 11 Commits
NLP		NLP
test		test
.gitignore		.gitignore
README.md		README.md
run_tests.py		run_tests.py

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

Custom NLP Library

Dependencies

Install Dependencies

Run Tests

About

Uh oh!

Releases

Packages

Uh oh!

Contributors

Uh oh!

Languages

Folders and files

Latest commit

History

Repository files navigation

Custom NLP Library

Dependencies

Install Dependencies

Run Tests

About

Topics

Resources

Uh oh!

Stars

Watchers

Forks

Releases

Packages 0

Uh oh!

Contributors

Uh oh!

Languages

Packages