Indian financial data processed for a Retrieval-Augmented Generation system. The dataset contains 24,780 text chunks with 384-dimensional embeddings generated by the all-MiniLM-L6-v2 model. It was created by user prakhar146 and last updated on April 11, 2026.
Use Cases
- Build a financial Q&A system based on the pre-indexed text chunks and embeddings.
- Retrieve relevant Indian market information based on the FAISS index for semantic search.
- Fine-tune or evaluate language models on Indian financial topics based on the provided text corpus.
Strengths
- Contains 24,780 pre-processed text chunks with embeddings, enabling immediate use for retrieval tasks.
- Includes a FAISS IndexFlatIP for fast similarity search, reducing implementation overhead.
- Chunks are processed with a 400-character size and 80-character overlap, which may help preserve context.
Limitations
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count for the original source data is unknown, which may limit suitability assessment.
- The description metadata is limited; actual data quality and coverage require manual inspection.
Provenance
- Source
- huggingface user prakhar146
- Collection Method
- Text chunks likely sourced from Indian financial documents or market data, processed and embedded.
- Time Range
- null
- Freshness
- Last updated 2026-04-11 05:09:04; freshness should be verified.
- Geography
- India, based on the dataset title and description of covering NSE/BSE markets.