Step 1: Image Capture
When you point your camera at a question or upload a photo, the first step is image preprocessing. The raw camera image goes through automatic adjustments for brightness, contrast, rotation, and perspective correction. If you photographed the page at a slight angle, the software straightens it. If the lighting was uneven, it normalizes the brightness across the image.
This preprocessing stage is critical because the accuracy of everything that follows depends on having a clean, well-exposed image. A blurry or dark photo will still be processed, but with lower confidence scores. The app warns you when image quality is too low for reliable results.
Step 2: Optical Character Recognition (OCR)
The preprocessed image is sent through our OCR engine, which is a specialized neural network trained to read text from images. Unlike general-purpose OCR systems (like those used for scanning receipts or business cards), our OCR is trained specifically on educational materials.
This means it understands mathematical notation (fractions, exponents, square roots, integrals, Greek letters), chemical formulas (subscripts, superscripts, reaction arrows, molecular structures), and scientific notation. It also handles multiple languages, different handwriting styles, and various fonts used by textbook publishers.
How OCR Handles Math
Mathematical OCR is significantly harder than text OCR. The spatial relationships between characters carry meaning. A "2" above and to the right of "x" means "x squared," while a "2" to the left of "x" means "2x." Our math OCR engine understands these spatial relationships and converts them into proper mathematical expressions that the AI can process.
Step 3: Question Understanding
Once the text is extracted, a natural language processing model analyzes what the question is actually asking. This is more complex than it sounds. The model needs to distinguish between an instruction ("solve for x"), a context sentence ("given the following data"), and the actual problem to solve.
For multi-part questions, it identifies each sub-question and understands dependencies between parts. If part (b) says "using your answer from part (a)," the model knows to solve part (a) first and use that result in part (b).
Step 4: AI Problem Solving
The understood question is passed to our AI reasoning engine. This is a large language model that has been fine-tuned on millions of solved educational problems across all subjects and grade levels. It does not look up answers in a database. It reasons through the problem the same way a knowledgeable tutor would.
For a math problem, it selects the appropriate method (factoring, quadratic formula, integration by parts), applies it step by step, and arrives at the answer. For a science question, it recalls the relevant laws and principles, applies them to the specific scenario, and generates an explanation. For humanities questions, it constructs a well-organized response using relevant facts and analysis.
Step 5: Answer Formatting
The raw AI output is formatted into a clean, readable answer. Mathematical expressions are rendered properly with correct notation. Steps are numbered and labeled. Key terms are highlighted. The final answer is clearly marked. This formatting makes it easy to read the solution and transcribe it to your homework.
Accuracy and Reliability
Our system achieves the following accuracy rates based on independent testing:
- Printed text recognition: 98.5% accuracy
- Handwritten text recognition: 91.2% accuracy
- Math problem solving: 96.1% accuracy
- Science question answering: 94.3% accuracy
- Humanities and language arts: 92.8% accuracy
These numbers represent the rate at which the complete solution is correct. When the system is unsure, it flags the answer with a confidence score so you know to double-check the solution.
Continuous Improvement and Model Updates
Our AI model is updated regularly to improve accuracy, add support for new question types, and incorporate advances in natural language processing. Unlike apps that require manual updates through the app store, our processing happens in the cloud. This means every improvement we make is instantly available to all users without any action on their part. When you scan a question today, you are always using the latest and most capable version of our AI engine.
The improvement process is driven by aggregate accuracy metrics. We track which types of questions produce the highest and lowest confidence scores across all users without identifying any individual user or their specific questions. When we identify a category of questions where accuracy is below our target threshold, our engineering team creates specialized training data for that category and fine-tunes the model. Recent improvements include better handling of organic chemistry nomenclature, improved accuracy on geometric proof problems, and enhanced support for questions written in non-English languages.
How We Handle Edge Cases
No AI system is perfect, and ours is designed to handle its limitations gracefully. When the OCR engine is not confident about a character or word it has read, it flags the uncertainty in the extracted text. The AI solver then considers multiple possible readings and selects the interpretation that makes the most sense in context. For example, if the OCR reads a character that could be either a "6" or a "b," the AI determines which interpretation produces a valid mathematical expression and uses that reading.
When the AI solver itself is uncertain about the answer, it provides a confidence score with the result. Scores above ninety percent indicate high reliability. Scores between seventy and ninety percent suggest the answer is likely correct but should be double-checked. Scores below seventy percent trigger a warning that the answer may not be reliable, and the system suggests rescanning with better image quality or consulting a human tutor for that specific problem. This transparent confidence scoring ensures you know exactly when to trust the answer and when additional verification is warranted.
Technology FAQ
What happens to my data after processing?
Your scanned images and generated answers are processed in real time and immediately discarded from our servers after delivery to your device. We do not store, analyze, or share your questions or answers. The processing pipeline is designed for privacy by default, meaning no personal academic data ever persists on our infrastructure. Aggregate performance metrics are collected without any connection to individual users or their specific questions, solely for the purpose of improving the AI model's accuracy across different question types.
How is this different from asking ChatGPT to solve a problem?
General-purpose chatbots like ChatGPT require you to type your question manually, which is impractical for complex equations, diagrams, and multi-part problems. Our system reads directly from photos, eliminating transcription errors and saving significant time. Additionally, our AI is specifically fine-tuned for educational problem solving, achieving higher accuracy on homework questions than general-purpose models. Our system also includes purpose-built OCR for mathematical notation and scientific formulas, confidence scoring for answer reliability, and batch processing for full worksheets, none of which are available through a general chatbot interface.