AI-Based Document Analysis and Question Answering System

Authors

  • Radhika Sharma Department of Electronics & Communication Engineering, Dr Akhilesh Das Gupta Institute of Professional Studies, GGSIPU, India
  • Devraj Gautam Department of Electronics & Communication Engineering, Dr Akhilesh Das Gupta Institute of Professional Studies, GGSIPU, India

DOI:

https://doi.org/10.65890/race.v2i2.192

Keywords:

Artificial Intelligence, Natural Language Processing, Amazon Web Services, Simple Storage Service, Application Programming Interface, Question Answering, Portable Document Format.

Abstract

The exponential growth of textual data, in the form of corporate documents, reports, and research papers, in the current digital era has increased the need for intelligent systems that can automatically comprehend documents. Analysing these documents by hand is ineffective and time-consuming. In order to extract valuable insights from documents, this study provides an AI- Based document analyzer with a question-answer system that makes use of Natural Language Processing approaches. Users can submit text or PDF files to the system, which then extracts content, generates succinct summaries, identifies keywords, and enables interactive question-answering. Python is used to build the architecture, which is then serverless deployed on Amazon Web Services (AWS) utilising Amazon S3 and Amazon EC2, and for the question-answering system, GEMINI is used. By reducing reading time and providing immediate access to pertinent information, the suggested solution increases productivity. It is affordable, scalable, and suitable for business, education, and research.

References

[1] S. Bird, E. Klein, and E. Loper, Natural Language Processing with Python, 1st ed. Sebastopol, CA, USA: O’Reilly Media, 2009.

[2] NLTK, “NLTK Documentation,” [Online]. Available: https://www.nltk.org/. [Accessed: Mar. 19, 2026].

[3] C. D. Manning, P. Raghavan, and H. Schütze, Introduction to Information Retrieval. Cambridge, U.K.: Cambridge University Press, 2008.

[4] T. Mikolov et al., “Efficient Estimation of Word Representations in Vector Space,” arXiv preprint arXiv:1301.3781, 2013

[5] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” in Proc. NAACL-HLT, 2019.

[6] PyPDF2, “PyPDF2 Documentation,” [Online]. Available: https://pypdf2.readthedocs.io/. [Accessed: Mar. 19, 2026].

[7] Amazon Web Services, “AWS Documentation,” [Online]. Available: https://docs.aws.amazon.com/. [Accessed: Mar. 19, 2026].

[8] Amazon Web Services, “Amazon S3 User Guide,” [Online]. Available: https://docs.aws.amazon.com/s3/. [Accessed: Mar. 19, 2026].

[9] Amazon Web Services, “AWS Lambda Developer Guide,” [Online]. Available: https://docs.aws.amazon.com/lambda/. [Accessed: Mar. 19, 2026].

[10] Amazon Web Services, “Amazon API Gateway Documentation,” [Online]. Available: https://docs.aws.amazon.com/apigateway/. [Accessed: Mar. 19, 2026].

[11] Napkin AI, “Napkin AI.” [Online]. Available: https://www.napkin.ai. [Accessed: Apr. 10, 2026].

Downloads

Published

31-07-2026

Data Availability Statement

The data used in this study are available from the corresponding author upon reasonable request. No publicly available dataset was used, and all experimental data were generated and analyzed during the current study.

Issue

Section

Research Articles

How to Cite

Sharma, R. ., & Gautam, D. (2026). AI-Based Document Analysis and Question Answering System. Revolutionary Advances in Computing and Electronics: An International Journal, 2(2), 17-32. https://doi.org/10.65890/race.v2i2.192